
Sep 9, 2026 · 40 min
Independent evaluations race to keep pace with AI models
Who Grades the AI Models? | Ben Horowitz & Rayan Krishnan
As models move into agents, enterprise workflows, and policy decisions, outdated or self-reported benchmarks can obscure both capability and risk.
- 1Independent evaluators must update benchmarks as models gain new capabilities and learn to optimize against existing tests.
- 2Enterprise model selection increasingly depends on task-specific evidence, pricing, tooling, and workflow performance rather than general rankings.
- 3Evaluations could give policymakers concrete evidence on alignment, cybersecurity, reward hacking, and geopolitical differences in model behavior.
Don't miss
Ryan Krishnan describes how VALS spent roughly $1.5 million in token value during one month using coding tools, leading to VALSmith’s cost-aware routing system.
The brief
Ben Horowitz and Ryan Krishnan open with a basic problem: public benchmarks can overstate what models do in practice, while labs’ self-reported results leave too much unexamined.
Krishnan explains VALS’s six-hour pre-release evaluations and its effort to turn fuzzy real-world abilities into explicit tests without slowing model launches or creating conflicts of interest.
The frontier is moving beyond question answering. VALS tests agents across long-running workflows and explores recursive improvement, where models help build improved versions of themselves.
For enterprises, the best general-purpose model may not fit a particular repository or task. VALSmith emerged after roughly $1.5 million in monthly token value, routing work to tools of different capability and cost.
The discussion ends with a policy question: can independent evaluations translate fast-changing model behavior into evidence that governments and international actors can use before standards fall behind?
Featuring
Listen to the full episode and explore every guest, topic, and moment on PodLume.

Ben Horowitz
Anthropic
OpenAI