Episode Details
Back to EpisodesBuilt to Fail: Why AI Benchmarks Have an Expiration Date
Season 1
Episode 293
Published 3 weeks, 2 days ago
Description
In five years, GPT-4 went from 27% on ARC (barely beating random guessing) to 96%. This isn't a success story—it's a cautionary tale about how benchmarks are designed to become obsolete. We trace the lifecycle of AI's most famous tests: born impossible, beaten quickly, replaced on schedule. Why do benchmarks die? Which ones survive? And what does it mean when a test designed to measure reasoning gets solved in a single model generation?
0:00 - The ARC Paradox: From Impossible to 96%
2:45 - What Is a Benchmark? (And Why You Only See One Number)
5:30 - The Canon: MMLU and the Benchmarks Everyone Quotes
11:00 - The Pattern: How Benchmarks Are Designed to Die
14:15 - Which Benchmarks Survive, and Why It Matters
---
Sources & further reading:
• All fetched and quoted on 2026-09-17.
• Hendrycks et al., Measuring Massive Multitask Language Understanding, arXiv:2009.03300
• 2020-09-07, ICLR 2021: https://arxiv.org/abs/2009.03300
• Zellers et al., HellaSwag, arXiv:1905.07830, 2019-05-19, ACL 2019
• Clark et al., Think you have Solved Question Answering? Try ARC, arXiv:1803.05457, 2018-03-14
• Cobbe et al., Training Verifiers to Solve Math Word Problems (GSM8K), arXiv:2110.14168
• 2021-10-27: https://arxiv.org/abs/2110.14168
• Hendrycks et al., Measuring Mathematical Problem Solving With the MATH Dataset
• arXiv:2103.03874, 2021-03-05, NeurIPS 2021: https://arxiv.org/abs/2103.03874
• Chen et al., Evaluating Large Language Models Trained on Code (HumanEval), arXiv:2107.03374
• 2021-07-07: https://arxiv.org/abs/2107.03374
• Rein et al., GPQA: A Graduate-Level Google-Proof Q&A Benchmark, arXiv:2311.12022, 2023-11-20
• the 546/448/198 subsets and their human numbers: https://arxiv.org/abs/2311.12022
• Srivastava et al., Beyond the Imitation Game (BIG-bench), arXiv:2206.04615, 2022-06-09
• Yue et al., MMMU, arXiv:2311.16502, 2023-11-27, CVPR 2024 Oral
• Jimenez et al., SWE-bench, arXiv:2310.06770, 2023-10-10, ICLR 2024
• OpenAI, GPT-4 Technical Report, arXiv:2303.08774, 2023-03-15 — the saturation table
• Wang et al., MMLU-Pro, arXiv:2406.01574, 2024-06-03 — the plateau quote
• Suzgun et al., Challenging BIG-Bench Tasks (BBH), arXiv:2210.09261, 2022-10-17; Kazemi et al.
• BIG-Bench Extra Hard*, arXiv:2502.19187, 2025-02-26
• Glazer et al., FrontierMath, arXiv:2411.04872, 2024-11-07; tier and holdout details at
• v2 error-correction note at: https://epoch.ai/frontiermath/tiers-1-4/about
• Chollet et al., ARC-AGI-2, arXiv:2505.11831, announced 2025-03-24
• composition and human calibration; ARC-AGI-3 launch: https://arcprize.org/arc-agi/2
• 2026-03-25 and the 2026-09-03 Astra post at: https://arcprize.org/blog/
• Phan, Gatti, Han et al., Humanity's Last Exam, arXiv:2501.14249, 2025-01-24, Nature 649
• 2026-01-28: https://agi.safe.ai/
• Ott, Barbosa-Silva, Blagec, Brauner, Samwald, *Mapping global dynamics of benchmark creation
• and saturation in artificial intelligence*, Nature Communications 13:6793, 2022-11-10
• Akhtar, Reuel, Soni et al., When AI Benchmarks Plateau, arXiv:2602.16763, 2026-02-18
• the expert-curation finding: https://arxiv.org/abs/2602.16763
• Bean, Kearns, Romanou et al., Measuring what Matters: Construct Validity in LLM Benchmarks
• arXiv:2511.04703, NeurIPS 2025 D&B: https://arxiv.org/abs/2511.04703
• Burnham, GPQA Diamond: what's left, Epoch AI, 2025-05-30
• [internal] data/series/ml-volleyball/ep-07-a-metric-is-a-procedure.md — the chance-floor
• callback — .
This podcast episode was fully generated by AI — research, script, voices, and production. Built with Claude, Piper TTS, and automated pipeline tooling.