Episode Details

Back to Episodes

The Benchmark Mirage: Why That 88% Means Nothing

Season 1 Episode 292 Published 3 weeks, 2 days ago
Description
You trained a language model for months, ran it on MMLU, and got 48.8%. But the published version of the same model? 63.4%. Nothing was wrong—the model was identical, the weights were identical, the data was identical. So what changed? Everything you think you know about benchmark scores is missing the crucial context. In this premiere episode of a new series on LLM evaluation, Host A and Host B dig into what a benchmark actually is, why the same model can produce wildly different results, and why someone telling you "this model scored 88%" is leaving out most of the important information. Timestamps: 00:00 - Opening: The 15-point mystery 04:30 - What actually changed between the two runs? 08:15 - The tool administering the test matters more than you think 12:00 - Announcing the evaluation series 15:30 - Full transparency: learning LLM benchmarking live --- Sources & further reading: • Hugging Face, "What's going on with the Open LLM Leaderboard?", published 2023-06-23 • Liang et al., Holistic Evaluation of Language Models (HELM), arXiv:2211.09110 • UK AI Security Institute, Inspect docs — and: https://inspect.aisi.org.uk/ • Hendrycks et al., Measuring Massive Multitask Language Understanding (MMLU) • Touvron et al., LLaMA: Open and Efficient Foundation Language Models • Wei et al., Chain-of-Thought Prompting Elicits Reasoning in Large Language Models • Wang et al., Self-Consistency Improves Chain of Thought Reasoning in Language Models • Sclar et al., *Quantifying Language Models' Sensitivity to Spurious Features in Prompt • EleutherAI, lm-evaluation-harness • Hugging Face, "Performances are plateauing, let's make the leaderboard steep again" • All URLs fetched and quoted from on 2026-09-17. • the three-implementation MMLU table: https://huggingface.co/blog/open-llm-leaderboard-mmlu • the prompt/scoring mechanics, the "not at all comparable" and "tied to their implementations" • quotes. Compared codebases: Eleuther harness commit e47e01b, HELM commit cab5d89, Hendrycks • hendrycks/test PR #13. • 2022-11-16, TMLR 2023 — — the (scenario, adaptation, metric): https://arxiv.org/abs/2211.09110 • triple; the OPT-175B HellaSwag 79.1 → 30.2 prompt-format finding. *Quotes taken from the ar5iv • HTML rendering of the arXiv source, not the TMLR PDF.* • (MIT licensed) — Task = Dataset + Solver +: https://github.com/UKGovernmentBEIS/inspect_ai • Scorer. *Note for ep-9/11: the docs homepage credits "the UK AI Security Institute and Meridian • Labs" while the GitHub README credits only AISI.* • arXiv:2009.03300, submitted 2020-09-07, ICLR 2021: https://arxiv.org/abs/2009.03300 • arXiv:2302.13971, Table 9 — — the published 63.4 five-shot MMLU: https://arxiv.org/abs/2302.13971 • figure for LLaMA 65B. • arXiv:2201.11903, 2022-01-28 — — PaLM 540B GSM8K 17.9 → 56.9.: https://arxiv.org/abs/2201.11903 • arXiv:2203.11171, 2022-03-21, ICLR 2023 — — GSM8K +17.9.: https://arxiv.org/abs/2203.11171 • Design...*, arXiv:2310.11324, 2023-10-17, ICLR 2024 — — up to: https://arxiv.org/abs/2310.11324 • 76 accuracy points across equivalent formats; ~10 average over 50+ tasks; median spread 7.5. • the README that never defines "harness".: https://github.com/EleutherAI/lm-evaluation-harness • Formal citation: Gao et al., The Language Model Evaluation Harness, Zenodo v0.4.3, July 2024 • doi:10.5281/zenodo.12608602. The v0.4.3 citation is NOT the version that produced the 0.488. • [preprint, peer review status unknown] Zhao, Wang, Bangash, Adams, Hassan, *Towards Evaluation • Engineering: An Empirical Study of ML Evaluation Harnesses in the Wild*, arXiv:2605.24213 • 2026-05-22 — — the nearest thing to a formal definition of: https://arxiv.org/abs/2605.24213 • "evaluation harness". • (Open LLM Leaderboard v2), published 2024-06-26 • saturation, contamination, benchmark: https://huggingface.co/spaces/open-llm-leaderboard/blog • …and 5 more in the episode research notes This podcast episode was fully generated by AI — research, script, voices, and production. Built w
Listen Now

Love PodBriefly?

If you like Podbriefly.com, please consider donating to support the ongoing development.

Support Us