
Sep 22, 2026 · 39 min
Claude and GPT-6 face a live blind test
Opus 5.5 vs. GPT-6 Sol: which model won my blind taste test?
The episode turns simultaneous model launches into a practical comparison of capability, workflow fit, and evaluator bias.
- 1Claire Vo replaces a prerecorded review with live coverage after Anthropic and OpenAI launch models the same morning.
- 2Her blind benchmark tests practical work across writing, product specifications, coding, prototypes, and creative tasks.
- 3The comparison is designed to expose both model strengths and the personal biases shaping AI evaluations.
Don't miss
Vo turns the benchmark inward, using the blind comparison to examine her own biases as an evaluator.
The brief
Claire Vo planned a prerecorded review of Claude Opus 5.5, but simultaneous launches from Anthropic and OpenAI pushed her into live coverage.
The episode frames model comparison as a practical benchmark rather than a leaderboard exercise, spanning writing, product specs, coding, prototypes, and creative work.
GPT-6 Sol and GPT-6 Luna enter the comparison alongside Claude, giving the test a broader question than which launch sounds most impressive.
Vo’s blind How I AI benchmark is meant to reveal where models excel while limiting the influence of brand expectations and personal preference.
The standout tension is methodological: a supposedly objective model test may also expose the evaluator’s own biases.
Featuring
Books & mentions
Listen to the full episode and explore every guest, topic, and moment on PodLume.

Claude
Anthropic
OpenAI
GPT-6 Sol
GPT-6 Luna