Episode Details
Back to EpisodesThe Confident Shrug: Why AI Can't Explain Itself
Season 1
Episode 332
Published 1 week, 5 days ago
Description
Modern AI can diagnose skin cancer with 95% accuracy—but ask why, and you get a confident shrug. We explore interpretability and explainability: why understanding what your AI models are actually thinking has become a legal requirement. From Grad-CAM visualizations to mechanistic interpretability, we map the techniques reshaping how we trust our models.
00:00:00 - The Doctor's Shrug
00:07:08 - Grad-CAM Visualization
00:10:44 - Mechanistic Interpretability
00:17:44 - The MRI vs Autopsy Framework
00:20:47 - What's Next
---
Sources & further reading:
• Ramprasaath R. Selvaraju et al. — "Grad-CAM: Visual Explanations from Deep Networks via Gradient-based Localization" (ICCV 2017) — [: https://arxiv.org/abs/1610.02391](https://arxiv.org/abs/1610.02391
• Marco Tulio Ribeiro, Sameer Singh, Carlos Guestrin — "'Why Should I Trust You?' Explaining the Predictions of Any Classifier" (KDD 2016) — [: https://arxiv.org/pdf/1602.04938](https://arxiv.org/pdf/1602.04938
• Andy Arditi, Oscar Obeso, Aaquib Syed, Daniel Paleka, Nina Panickssery, Wes Gurnee — "Refusal in Language Models Is Mediated by a Single Direction" (NeurIPS 2024) — [: https://arxiv.org/abs/2406.11717](https://arxiv.org/abs/2406.11717
• Anthropic Research — "On the Biology of a Large Language Model" (Circuit Tracing, March 2025) — [: https://transformer-circuits.pub/2025/attribution-graphs/biology.html](https://transformer-circuits.pub/2025/attribution-graphs/biology.html
• Anthropic — Interpretability Research page — [: https://www.anthropic.com/research/team/interpretability](https://www.anthropic.com/research/team/interpretability
• DeepSci — "Mechanistic Interpretability in 2026: Can We See Inside AI Models?" — [: https://deepsci.io/blog/mechanistic-interpretability-2026](https://deepsci.io/blog/mechanistic-interpretability-2026
• Towards AI / Yuval Mehta — "Mechanistic Interpretability Is Having Its Moment" — [: https://pub.towardsai.net/mechanistic-interpretability-is-having-its-moment-what-engineers-actually-need-to-know-e4421f305f84](https://pub.towardsai.net/mechanistic-interpretability-is-having-its-moment-what-engineers-actually-need-to-know-e4421f305f84
• Subhadip Mitra — "Circuit Tracing for the Rest of Us" — [: https://subhadipmitra.com/blog/2026/circuit-tracing-production/](https://subhadipmitra.com/blog/2026/circuit-tracing-production/
• Salih et al. — "A Perspective on Explainable Artificial Intelligence Methods: SHAP and LIME" (Advanced Intelligent Systems, 2025) — [: https://advanced.onlinelibrary.wiley.com/doi/10.1002/aisy.202400304](https://advanced.onlinelibrary.wiley.com/doi/10.1002/aisy.202400304
• Nature Scientific Reports — "A comparative evaluation of explainability techniques for image data" (2025) — [: https://www.nature.com/articles/s41598-025-25839-y](https://www.nature.com/articles/s41598-025-25839-y
• Human Rights Pulse — "Dutch court finds SyRI algorithm violates human rights norms" — [: https://www.humanrightspulse.com/mastercontentblog/dutch-court-finds-syri-algorithm-violates-human-rights-norms-in-landmark-case](https://www.humanrightspulse.com/mastercontentblog/dutch-court-finds-syri-algorithm-violates-human-rights-norms-in-landmark-case
• Library of Congress — "Netherlands: Court Prohibits Government's Use of AI Software to Detect Welfare Fraud" — [: https://www.loc.gov/item/global-legal-monitor/2020-03-13/netherlands-court-prohibits-governments-use-of-ai-software-to-detect-welfare-fraud/](https://www.loc.gov/item/global-legal-monitor/2020-03-13/netherlands-court-prohibits-governments-use-of-ai-software-to-detect-welfare-fraud/
• …and 6 more in the episode research notes
This podcast episode was fully generated by AI — research, script, voices, and production. Built with Claude, Piper TTS, and automated pipeline tooling.