vLLM
Product
vLLM is an open-source, high-throughput inference and serving engine for large language models, designed for efficient production deployment.
4 episodes featuring vLLM

TBPN
Meta bets personal AI can make glasses everyday computers
The episode connects Meta’s consumer-AI strategy to the infrastructure, reliability, security, and planetary-engineering bets shaping technology’s next phase.
Sep 24, 2026 · 2h 2m

Inside AsembleAI: DeepTech, AI & Science
Enterprise AI needs a control plane for models and costs
As organizations adopt more specialized and open-weight models, governance, data protection, and inference economics become operational necessities rather than optional safeguards.
Sep 15, 2026 · 22 min

TBPN
AI risk meets robot arms, agent investing and chip competition
The episode connects abstract arguments about AI safety to hands-on robotics, emerging agent businesses and the hardware race challenging NVIDIA.
Sep 11, 2026 · 1h 46m

The a16z Show
Open models turn inference into critical AI infrastructure
The episode examines whether open-weight models can match closed systems while giving organizations more control over cost, deployment, customization, and safety.
Aug 6, 2026 · 47 min
