Episode Details
Back to EpisodesThe Benchmark Is Lying (But Not How You Think)
Season 1
Episode 296
Published 3 weeks, 2 days ago
Description
When the same AI model scores 18.33% and 38% on the identical benchmark with no weight changes, what actually changed? Spoiler: the scaffolding. This episode dismantles the illusion of model leaderboards by revealing they measure entire systems—model plus prompt plus tools plus sampling—not the model in isolation. We break down three separate knobs you can turn to swing scores by double digits, examine why the labs already know this and still publish numbers anyway, and explore what a real model comparison might look like.
00:00 - The Setup: 18.33% vs 38%
02:30 - What Changed? Only the Packaging
05:45 - The Core Thesis: You're Benchmarking a System
08:15 - The Anthropic Blog: The Labs Are Already Telling You This
11:20 - Three Mechanisms That Move the Needle
15:00 - What This Means for AI Credibility
18:30 - Outro
---
Sources & further reading:
• All fetched and quoted on 2026-09-17.
• Anthropic, Raising the bar on SWE-bench Verified / engineering post
• live page reads "Published Jan 06: https://www.anthropic.com/engineering/swe-bench-sonnet
• 2025"; same content first published late October 2024 at /research/swe-bench-sonnet. The
• "entire agent system" and "can vary significantly based on this scaffolding" quotes.
• Xia, Deng, Dunn, Zhang, Agentless, arXiv:2407.01489, 2024-07-01
• Table 1, the GPT-4o and GPT-4 per-scaffold spreads with costs: https://arxiv.org/abs/2407.01489
• and token counts.
• Anthropic, Claude 3.7 Sonnet and Claude Code, 2025-02-24
• 63.7% → 70.3% same model; the 489/500 subset.: https://www.anthropic.com/news/claude-3-7-sonnet
• Anthropic, Claude 4, 2025-05-22 — — 72.5/72.7 → 79.4/80.2: https://www.anthropic.com/news/claude-4
• with parallel attempts and a scoring model; the dropped third planning tool.
• Anthropic, Claude's extended thinking, 2025-02-24
• "the very same model… more time": https://www.anthropic.com/news/visible-extended-thinking
• GPQA 84.8% at 256 samples + 64k thinking.
• Brown, Juravsky, Ehrlich, Clark, Le, Ré, Mirhoseini, Large Language Monkeys
• arXiv:2407.21787, 2024-07-31 — 15.9% → 56% on SWE-bench Lite at fixed weights, and the
• verifier-free plateau.
• Snell, Lee, Xu, Kumar, Scaling LLM Test-Time Compute Optimally…, arXiv:2408.03314, 2024-08-06
• cited with the caveat that its mechanisms include a trained process verifier and a revision
• model, so it is not a pure frozen-weights result.
• Gao, Madaan, Zhou, Alon, Liu, Yang, Callan, Neubig, PAL: Program-aided Language Models
• arXiv:2211.10435, 2022-11-18 — 19.7 → 65.6 → 72.0 → 80.4 on one Codex checkpoint, and the
• brittleness contrast (65.6 → 20.1 vs 72.0 → 61.5).
• METR, Measuring AI Ability to Complete Long Software Tasks, arXiv:2503.14499 §E.4 — the
• "very large difference" and "reasonable lower bound" quotes, and the 2–3 engineer weeks asymmetry.
• METR, Guidelines for capability elicitation, 2024-03-15
• the minimum scaffold: https://metr.org/blog/2024-03-15-guidelines-for-capability-elicitation/
• and the spurious-failure taxonomy.
• UK AI Safety/Security Institute, Advanced AI evaluations: May update, 2024-05-20
• their scaffold at 25% vs 24%: https://www.aisi.gov.uk/blog/advanced-ai-evaluations-may-update
• and 32%.
• [preprint, unreviewed] Harness-Bench, arXiv:2605.27922, 2026-05-27 — 6 harnesses × 8 backends
• 5,088 trajectories, the 23.8-point gap, and the configuration-level reporting recommendation.
• [preprint, unreviewed position paper] Stop Comparing LLM Agents Without Disclosing the Harness
• arXiv:2605.23950, 2026-05-07 — the Binding Constraint Thesis and the ranking-reversal claim, with
• its comparable-frontier-capability scope.
• Biderman, Schoelkopf, Sutawika, Gao et al., *Lessons from the Trenches on Reproducible
• Evaluation of Language Models*, arXiv:2405.14782 — the reporting best practices and the MMLU
• micro-vs-macro averaging point.
• …and 8 more in the episode research notes
This podcast episode was fully generated by AI — research, script, voices, and production. Built with Claude, Piper TT