Episode Details

Back to Episodes

The Blank Exam That Got an A: When AI Judges Can't Judge

Season 1 Episode 295 Published 3 weeks, 2 days ago
Description
When a constant response beats cutting-edge AI models on major benchmarks, something's profoundly broken with how we evaluate AI progress. Researchers built a 'no-model' that scored 86.5% on Alpaca by exploiting formatting tricks instead of reading answers, revealing fundamental flaws in three major AI benchmarks. The hosts explore how evaluation systems get gamed and why scoring turned out to be the hardest part. 00:00 - The Paradox: An Answer That's Not An Answer 03:00 - The No-Model Research and Benchmark Results 08:00 - How It Works: Gaming the Judge's Evaluation 14:00 - Defeating Benchmark Defenses 20:00 - Outro --- Sources & further reading: • All fetched and quoted on 2026-09-17. • Zheng, Chiang, Sheng, Zhuang et al., *Judging LLM-as-a-Judge with MT-Bench and Chatbot • Arena*, arXiv:2306.05685, 2023-06-09, NeurIPS 2023 D&B: https://arxiv.org/abs/2306.05685 • agreement numbers, the four biases, and the §D.3 admission about the human baseline. Quotes • verified against the published NeurIPS PDF. • Wang, Li, Chen et al., Large Language Models are not Fair Evaluators, arXiv:2305.17926 • 2023-05-29 — the 66-of-80 position flip and the conflict-rate-by-quality-gap table. • Panickssery, Bowman, Feng, LLM Evaluators Recognize and Favor Their Own Generations • arXiv:2404.13076, 2024-04-15 — self-recognition 73.5%, >90% fine-tuned, and the correlation • with self-preference. • Zheng, Pang, Du et al., *Cheating Automatic LLM Benchmarks: Null Models Achieve High Win • Rates*, arXiv:2410.07137, 2024-10-09, ICLR 2025 Oral — 86.5% / 83.0 / 9.55, and the • swap-resilient structure. • Chen et al., Evaluating Large Language Models Trained on Code, arXiv:2107.03374 — the pass@k • estimator, the biased shortcut, and the per-k optimal temperatures. • OpenAI, HealthBench, arXiv:2505.08775, 2025-05-13 — 262 physicians, 48,562 criteria • grader macro-F1, and the 55–75% agreement ceiling. • OpenAI, Introducing SWE-bench Verified, 2024-08-13 — 61.1% unfair tests, 68.3% filtered • GPT-4o 16% → 33.2%; and *Why SWE-bench Verified no longer measures frontier coding • capabilities*, 2026-02-23 — the 59.4% figure. • UK AISI Inspect model-graded scorers: https://inspect.aisi.org.uk/model-graded.html • GRADE: C/GRADE: I extraction, last-grade binding, delimiter neutralisation, and the caveat • that graders run at non-zero temperature with no fixed seed, so borderline grades move between • runs. Docs are unversioned; cite as accessed 2026-09-17. • inspectevals issues #2292 and #2293 and PR #2294, UKGovernmentBEIS/inspectevals • opened 2026-08-25, all open and unacknowledged by maintainers as of 2026-09-17 — verified via • the GitHub API and by reading the classifier source on main. • Cohen 1960, EPM 20(1):37–46, doi:10.1177/001316446002000104 · Fleiss 1971, Psychological • Bulletin 76(5):378–382, doi:10.1037/h0031619 · Krippendorff 1970, EPM 30(1):61–70 • doi:10.1177/001316447003000105 · Landis & Koch 1977, Biometrics 33(1):159–174 • doi:10.2307/2529310 (pagination is 159–174; the widely-circulated 150–174 is wrong). Interiors • of these three are paywalled — the "clearly arbitrary" quote comes via Löwe, *Measuring the • Agreement of Mathematical Peer Reviewers*, Global Philosophy, 2022-12-21 • doi:10.1007/s10516-022-09647-x, which cites Landis & Koch p. 164. • Zapf, Castell, Morawietz, Karch, Measuring inter-rater reliability for nominal data, BMC • Medical Research Methodology, 2016 — the Fleiss/Scott's-pi misnomer, missing-data behaviour, and • the prevalence objection to fixed cut-offs. • [internal] GetTheJob/research/gamesenser-technical-profile.md — 19 matches / 721 rallies • 4 matches / 755 touches, ~50 min marking per set, blind panels with sealed keys, and the • BULK-1 box-score validation including the 627% line reported with n = 9 — . This podcast episode was fully generated by AI — research, script, voices, and production. Built with Claude, Piper TTS, and automated pipeline tooling.
Listen Now

Love PodBriefly?

If you like Podbriefly.com, please consider donating to support the ongoing development.

Support Us