IP

Inference pipeline parallelism

Topic

What experts have said about Inference pipeline parallelism

2 statements · 2 neutral

  1. Reiner PopeNeutralApr 29, 2026· Dwarkesh Podcast

    Inference pipelining does not materially reduce memory time or compute time.

    in inference. What are we saving on? Are we saving on memory time or compute time? Not really.

    Listen at 55:35

    Open the episode · Reiner Pope – The math behind how LLMs are trained and served
  2. Reiner PopeNeutralApr 29, 2026· Dwarkesh Podcast

    Inference pipelining is neutral for batch size and latency.

    inference, actually the effect of pipelining on anything you care about like batch size or latency actually is neutral

    Listen at 1:00:59

    Open the episode · Reiner Pope – The math behind how LLMs are trained and served

Statements are attributed to the speaker as said on the episode and reflect their view at the time, not PodLume's. They are not advice.

What is PodLume?

PodLume turns podcasts into searchable knowledge. AI-decoded transcripts, identified guests and topics, smart highlights, and cross-show search across the world’s best conversations — all in your pocket.

Inference pipeline parallelism | PodLume