How I AI
How I AI

Sep 22, 2026 · 39 min

Claude and GPT-6 face a live blind test

Opus 5.5 vs. GPT-6 Sol: which model won my blind taste test?

The episode turns simultaneous model launches into a practical comparison of capability, workflow fit, and evaluator bias.

3 key takeaways
  1. 1Claire Vo replaces a prerecorded review with live coverage after Anthropic and OpenAI launch models the same morning.
  2. 2Her blind benchmark tests practical work across writing, product specifications, coding, prototypes, and creative tasks.
  3. 3The comparison is designed to expose both model strengths and the personal biases shaping AI evaluations.

Don't miss

Vo turns the benchmark inward, using the blind comparison to examine her own biases as an evaluator.

The brief

Claire Vo planned a prerecorded review of Claude Opus 5.5, but simultaneous launches from Anthropic and OpenAI pushed her into live coverage.

The episode frames model comparison as a practical benchmark rather than a leaderboard exercise, spanning writing, product specs, coding, prototypes, and creative work.

GPT-6 Sol and GPT-6 Luna enter the comparison alongside Claude, giving the test a broader question than which launch sounds most impressive.

Vo’s blind How I AI benchmark is meant to reveal where models excel while limiting the influence of brand expectations and personal preference.

The standout tension is methodological: a supposedly objective model test may also expose the evaluator’s own biases.

Featuring

Books & mentions

Listen to the full episode and explore every guest, topic, and moment on PodLume.

What is PodLume?

PodLume turns podcasts into searchable knowledge. AI-decoded transcripts, identified guests and topics, smart highlights, and cross-show search across the world’s best conversations — all in your pocket.

Claude and GPT-6 face a live blind test | PodLume