Inside AsembleAI: DeepTech, AI & Science

Diffusion models target real-time voice intelligence

EP 58: Every Millisecond Matters: Diffusion LLMs and the Future of Voice AI | Aditya Grover, Inception

Voice agents must respond quickly without sacrificing the reasoning quality that makes them useful in complex, multi-step interactions.

3 key takeaways
  1. 1Latency compounds across voice turns, agent steps, and tool calls, making milliseconds a core product constraint.
  2. 2Mercury generates and refines complete responses in parallel instead of producing tokens sequentially, aiming to preserve quality at conversational speed.
  3. 3Inception frames the next phase of AI around value per dollar and watt, not simply larger models or higher token counts.

Don't miss

Aditya Agarwal explains how Mercury begins with a rough complete response and refines it in parallel to reduce latency without abandoning quality.

The brief

Inception co-founder and CTO Aditya Agarwal argues that latency is no longer a technical footnote: every voice turn, agent step, and tool call compounds delay.

Mercury takes a different route from conventional language models, starting with a rough full-answer guess and refining it in parallel rather than emitting one token at a time.

The central tradeoff is speed versus intelligence. Inception is positioning Mercury for voice agents that respond in real time while retaining the quality expected from newer reasoning systems.

Agarwal says Mercury is designed to fit existing production stacks through an OpenAI-compatible API, while open-source release is not on the near-term roadmap.

The broader thesis is “value maxing”: as companies measure throughput, cost, and energy, intelligence per dollar and watt may matter more than maximizing tokens.

Agarwal expects voice agents to become smarter, more natural, and less interrupted over the next five to ten years, especially in customer support and other intelligent interactions.

Listen to the full episode and explore every guest, topic, and moment on PodLume.

What is PodLume?

PodLume turns podcasts into searchable knowledge. AI-decoded transcripts, identified guests and topics, smart highlights, and cross-show search across the world’s best conversations — all in your pocket.

Diffusion models target real-time voice intelligence | PodLume