LA
Long-context attention compute
Topic
What experts have said about Long-context attention compute
1 statement · 1 negative
Attention’s context-dependent compute cost becomes noticeable at contexts of millions of tokens.
“You start to notice the effect of the quadratic or the linear term up in the millions of tokens or so.”
Open the episode · Reiner Pope – The math behind how LLMs are trained and servedListen at 1:52:24
Statements are attributed to the speaker as said on the episode and reflect their view at the time, not PodLume's. They are not advice.
