Episode Details

Back to Episodes

Built to Fail: Why AI Benchmarks Have an Expiration Date

Season 1 Episode 293 Published 3 weeks, 2 days ago
Description
In five years, GPT-4 went from 27% on ARC (barely beating random guessing) to 96%. This isn't a success story—it's a cautionary tale about how benchmarks are designed to become obsolete. We trace the lifecycle of AI's most famous tests: born impossible, beaten quickly, replaced on schedule. Why do benchmarks die? Which ones survive? And what does it mean when a test designed to measure reasoning gets solved in a single model generation? 0:00 - The ARC Paradox: From Impossible to 96% 2:45 - What Is a Benchmark? (And Why You Only See One Number) 5:30 - The Canon: MMLU and the Benchmarks Everyone Quotes 11:00 - The Pattern: How Benchmarks Are Designed to Die 14:15 - Which Benchmarks Survive, and Why It Matters --- Sources & further reading: • All fetched and quoted on 2026-09-17. • Hendrycks et al., Measuring Massive Multitask Language Understanding, arXiv:2009.03300 • 2020-09-07, ICLR 2021: https://arxiv.org/abs/2009.03300 • Zellers et al., HellaSwag, arXiv:1905.07830, 2019-05-19, ACL 2019 • Clark et al., Think you have Solved Question Answering? Try ARC, arXiv:1803.05457, 2018-03-14 • Cobbe et al., Training Verifiers to Solve Math Word Problems (GSM8K), arXiv:2110.14168 • 2021-10-27: https://arxiv.org/abs/2110.14168 • Hendrycks et al., Measuring Mathematical Problem Solving With the MATH Dataset • arXiv:2103.03874, 2021-03-05, NeurIPS 2021: https://arxiv.org/abs/2103.03874 • Chen et al., Evaluating Large Language Models Trained on Code (HumanEval), arXiv:2107.03374 • 2021-07-07: https://arxiv.org/abs/2107.03374 • Rein et al., GPQA: A Graduate-Level Google-Proof Q&A Benchmark, arXiv:2311.12022, 2023-11-20 • the 546/448/198 subsets and their human numbers: https://arxiv.org/abs/2311.12022 • Srivastava et al., Beyond the Imitation Game (BIG-bench), arXiv:2206.04615, 2022-06-09 • Yue et al., MMMU, arXiv:2311.16502, 2023-11-27, CVPR 2024 Oral • Jimenez et al., SWE-bench, arXiv:2310.06770, 2023-10-10, ICLR 2024 • OpenAI, GPT-4 Technical Report, arXiv:2303.08774, 2023-03-15 — the saturation table • Wang et al., MMLU-Pro, arXiv:2406.01574, 2024-06-03 — the plateau quote • Suzgun et al., Challenging BIG-Bench Tasks (BBH), arXiv:2210.09261, 2022-10-17; Kazemi et al. • BIG-Bench Extra Hard*, arXiv:2502.19187, 2025-02-26 • Glazer et al., FrontierMath, arXiv:2411.04872, 2024-11-07; tier and holdout details at • v2 error-correction note at: https://epoch.ai/frontiermath/tiers-1-4/about • Chollet et al., ARC-AGI-2, arXiv:2505.11831, announced 2025-03-24 • composition and human calibration; ARC-AGI-3 launch: https://arcprize.org/arc-agi/2 • 2026-03-25 and the 2026-09-03 Astra post at: https://arcprize.org/blog/ • Phan, Gatti, Han et al., Humanity's Last Exam, arXiv:2501.14249, 2025-01-24, Nature 649 • 2026-01-28: https://agi.safe.ai/ • Ott, Barbosa-Silva, Blagec, Brauner, Samwald, *Mapping global dynamics of benchmark creation • and saturation in artificial intelligence*, Nature Communications 13:6793, 2022-11-10 • Akhtar, Reuel, Soni et al., When AI Benchmarks Plateau, arXiv:2602.16763, 2026-02-18 • the expert-curation finding: https://arxiv.org/abs/2602.16763 • Bean, Kearns, Romanou et al., Measuring what Matters: Construct Validity in LLM Benchmarks • arXiv:2511.04703, NeurIPS 2025 D&B: https://arxiv.org/abs/2511.04703 • Burnham, GPQA Diamond: what's left, Epoch AI, 2025-05-30 • [internal] data/series/ml-volleyball/ep-07-a-metric-is-a-procedure.md — the chance-floor • callback — . This podcast episode was fully generated by AI — research, script, voices, and production. Built with Claude, Piper TTS, and automated pipeline tooling.
Listen Now

Love PodBriefly?

If you like Podbriefly.com, please consider donating to support the ongoing development.

Support Us