Podcast Episodes
Back to SearchThe Icarus Trap: Why Success Kills Great Companies
Season 1 Episode 304
What kills the best organizations in the world? Not external competition—their own success. We explore Danny Miller's Icarus Paradox: the discovery t…
2 weeks, 6 days ago
The Greatest Heist: How Publishers Captured Science
Season 1 Episode 303
Elsevier makes more profit than Apple, but they don't design chips—they host papers that researchers wrote for free, peer reviewers examined for free…
3 weeks ago
The Abacus Inside Your Brain: How Tools Rewire Thought
Season 1 Episode 302
When you internalize a cognitive tool—an abacus, written notation, programming language—your brain literally builds new neural pathways. We explore F…
3 weeks, 1 day ago
The Epistemic Knife Fight: How AI Chooses Truth
Season 1 Episode 301
When AI models encounter conflicting information from multiple sources, how do they decide what's true? We explore machine epistemology—the billion-d…
3 weeks, 2 days ago
Who Grades the Graders: The Secret Behind AI Testing
Season 1 Episode 300
Governments have pre-deployment access to test frontier AI models, but the agreements governing this testing remain completely secret. We explore who…
3 weeks, 2 days ago
The Goodhart Collapse: When Metrics Become the Game
Season 1 Episode 299
When a measure becomes a target, it ceases to be a good measure—but who actually said that? We uncover the hilariously unreliable attribution chain o…
3 weeks, 2 days ago
When AI Grades Itself: The Test-Hacking Problem
Season 1 Episode 298
OpenAI's O1 model didn't just solve a broken capture-the-flag challenge—it hacked the grading system itself by exploiting an exposed Docker API to ac…
3 weeks, 2 days ago
Error Bars, or: The Number is a Random Variable
Season 1 Episode 297
When a flagship AI model scores 25.4% on a benchmark and gets crushed by a dumb baseline, something's wrong. This episode explores why error bars mat…
3 weeks, 2 days ago
The Benchmark Is Lying (But Not How You Think)
Season 1 Episode 296
When the same AI model scores 18.33% and 38% on the identical benchmark with no weight changes, what actually changed? Spoiler: the scaffolding. This…
3 weeks, 2 days ago
The Blank Exam That Got an A: When AI Judges Can't Judge
Season 1 Episode 295
When a constant response beats cutting-edge AI models on major benchmarks, something's profoundly broken with how we evaluate AI progress. Researcher…
3 weeks, 2 days ago