
Aug 6, 2026 · 47 min
Open models turn inference into critical AI infrastructure
The Engine Powering Open-Source AI
The episode examines whether open-weight models can match closed systems while giving organizations more control over cost, deployment, customization, and safety.
- 1vLLM evolved from an academic project into infrastructure that makes open-weight models practical to deploy at scale.
- 2Open models are narrowing the capability gap, but training costs and restrictive licenses complicate the economics of openness.
- 3Distributed development can improve safety and innovation by letting organizations adapt models instead of relying solely on centralized APIs.
Don't miss
The speakers discuss how Hugging Face used an open model to help contain a cyberattack involving a proprietary model.
The brief
Simon Mou traces vLLM’s path from an academic project to an inference engine that turns GPUs and other accelerators into reliable model endpoints.
As models grew from BERT-era experiments into systems embedded in daily work, serving variable workloads efficiently became a distinct engineering and organizational challenge.
The case for open weights is practical as much as ideological: organizations want lower costs, infrastructure control, customization, and the ability to set their own guardrails.
The tension is that open models still carry enormous training and operational costs, pushing labs toward licenses that preserve access while seeking sustainable funding.
A discussion of Hugging Face using an open model to help contain a cyberattack illustrates the broader claim: distributed access can support safety, not merely weaken it.
Simon argues that open and closed models may soon differ more in distribution and business strategy than in capability, creating a shared global innovation race.
Featuring
Listen to the full episode and explore every guest, topic, and moment on PodLume.

Matt Bornstein
Hugging Face
Nvidia Corporation
OpenAI