Episode Details
Back to EpisodesThe Goodhart Collapse: When Metrics Become the Game
Season 1
Episode 299
Published 3 weeks, 2 days ago
Description
When a measure becomes a target, it ceases to be a good measure—but who actually said that? We uncover the hilariously unreliable attribution chain of Goodhart's Law (the quote about measurement reliability is itself unreliably sourced), then explore what happens when you add money and reputation to metrics, turning them into leaderboards where gaming the system becomes the real game.
00:00 - Opening: The Perfect Irony
03:00 - The Attribution Chain: Goodhart, Hoskin, Strathern
08:00 - What Goodhart Actually Observed (1975)
12:00 - The Benchmarking Series So Far
14:30 - When Leaderboards Add Incentives
19:00 - Gaming the System: Money, Reputation, and Collapse
---
Sources & further reading:
• All fetched and quoted on 2026-09-17.
• Strathern, 'Improving ratings': audit in the British University system, European Review
• 5(3):305–321, 1997, p. 308 — the famous sentence, the Hoskin attribution, and "measurement and
• target rise together". Read via a scan of the published article.
• Goodhart's 1975 formulation, quoted via Mattson, Bushardt, Artino, Journal of Graduate
• Medical Education 13(1):2–5, 2021-02-13, doi:10.4300/JGME-D-20-01492.1 — the 1975 RBA text itself
• could not be fetched, and the editorial cites two candidate papers.
• Schaeffer, Pretraining on the Test Set Is All You Need, arXiv:2309.08632, 2023-09-13
• phi-CTNL, 100% estimated contamination, and the satire disclaimer.
• OpenAI, GPT-4 Technical Report, arXiv:2303.08774, 2023-03-15 — the BIG-bench admission
• Table 11 contamination rates, and the GSM-8K "in-between" sentence.
• Zhang et al. (Scale AI), arXiv:2405.00332 — checkpoint selection as overfitting without
• contamination.
• ARC Prize, OpenAI o3 Breakthrough High Score on ARC-AGI-Pub, 2024-12-20, with later updates
• 75.7% / 87.5%, trained on 75% of the public: https://arcprize.org/blog/oai-o3-pub-breakthrough
• training set, and the 2025-04-16 confirmation that the shipped o3 differs from the tested one.
• Chiang et al., Chatbot Arena, arXiv:2403.04132, 2024-03-07 — Bradley-Terry rather than Elo.
• Operator history: (2024-03-01).: https://lmsys.org/blog/2024-03-01-policy/
• **Singh, Nan, Wang, D'Souza, Kapoor, Üstün, Koyejo, Deng, Longpre, Smith, Ermis, Fadaee
• Hooker, The Leaderboard Illusion*, arXiv:2504.20879, v1 2025-04-29 / v2 2025-05-12 — the 27
• Meta variants, the data-share figures, 205 silently removed models, and the 112%-on-ArenaHard
• figure (quote the v2 wording).
• LMArena / Arena Intelligence, Our Response to "The Leaderboard Illusion" Writeup
• 2025-05-09 — — the three concessions and every rebuttal: https://arena.ai/blog/our-response/
• quoted above.
• Meta AI, The Llama 4 herd, 2025-04-05: https://ai.meta.com/blog/llama-4-multimodal-intelligence/
• — the 1417 ELO for an unreleased experimental version. LMArena's objection is quoted as reported
• (simonwillison.net, 2025-04-08); the primary X post could not be fetched.
• Gemini Team, arXiv:2312.11805, and Google's Introducing Gemini blog, 2023-12-06 — 90.04%
• CoT@32 vs 83.7% 5-shot, GPT-4 at 87.29% under the same scheme, and the per-model benefit of the
• uncertainty-routed decoding.
• Anthropic, Claude 3.7 Sonnet and Claude Code, 2025-02-24 — the disclosed scaffold, the
• 63.7→70.3 delta, and the 489/500 denominator counted as failures.
• Epoch AI, Clarifying the creation and use of the FrontierMath benchmark, 2025-01-23
• the funding partnership, ownership, embargo: https://epoch.ai/latest/openai-and-frontiermath
• and the holdout still being finalised in January 2025.
• Epoch AI, Transparency — checked 2026-09-17 and: https://epoch.ai/about/transparency
• live, contradicting the 404 recorded on 2026-09-03. Donations ≥ $70,000 including Coefficient
• Giving, Jaan Tallinn, SFF, Schmidt Sciences and Leopold Aschenbrenner; the OpenAI / Google
• …and 8 more in the episode research notes
This podcast episode was fully generated by AI — research, script, voices, and production. Built with Claude, Piper TTS, and automated pipeline toolin