Episode Details
Back to EpisodesGrading the Grader: When Tests Fail to See
Season 1
Episode 323
Published 2 weeks, 1 day ago
Description
How good is your test suite really? Host A plants bugs while Host B watches tests pass anyway—some changes harmless, others cutting critical safety systems. They explore the mutation testing framework, kill rates, and the uncomfortable truth: a test suite that passes everything might be hiding catastrophic blind spots.
00:00 - Introduction: The bug-planting experiment
01:30 - First invisible bug: The harmless edit no test catches
05:00 - Second invisible bug: Removed backup interlock
08:00 - The paradox: Two invisible bugs, opposite meanings
10:00 - What is grading the grader? Mutation testing framework
12:00 - Episode 8 callback: The Compilation Illusion
14:00 - Kill rate scoring and the 90% benchmark
---
Sources & further reading:
• [peer-reviewed] Jia and Harman, *An Analysis and Survey of the Development of Mutation
• Testing*, IEEE TSE 37(5):649–678, 2011, doi:10.1109/TSE.2010.62 — authors' copy
• read 2026-09-24: http://crest.cs.ucl.ac.uk/fileadmin/crest/sebasepaper/JiaH10.pdf
• (killed/survived definition, equivalent mutants, undecidability, mutation-score definition
• quoted verbatim above).
• [preprint] Evan Miller, *Adding Error Bars to Evals: A Statistical Approach to Language Model
• Evaluations*, arXiv:2411.00640 — — checked 2026-09-24: https://arxiv.org/abs/2411.00640
• (quote and submission date).
This podcast episode was fully generated by AI — research, script, voices, and production. Built with Claude, Piper TTS, and automated pipeline tooling.