LI
LLM inference queueing
Topic
What experts have said about LLM inference queueing
1 statement · 1 negative
A 20-millisecond inference schedule can produce up to 40 milliseconds of worst-case latency.
“the worst case latency is 40 milliseconds”
Open the episode · Reiner Pope – The math behind how LLMs are trained and servedListen at 23:35
Statements are attributed to the speaker as said on the episode and reflect their view at the time, not PodLume's. They are not advice.
