LI
LLM inference latency
Topic
What experts have said about LLM inference latency
1 statement · 1 negative
Inference latency has a lower bound set by reading all model parameters from memory.
“there is a lower bound on latency, which is simply. Simply I need to read all of my total parameters from memory into the chips”
Open the episode · Reiner Pope – The math behind how LLMs are trained and servedListen at 10:28
Statements are attributed to the speaker as said on the episode and reflect their view at the time, not PodLume's. They are not advice.
