Episode Details

Back to Episodes

Prions and Poison: When 0.001% Breaks Everything

Season 1 Episode 290 Published 3 weeks, 2 days ago
Description
What happens when someone intentionally poisons an AI's training data to make it dangerously wrong? Host A and Host B explore how researchers showed that replacing just 0.001% of training tokens with misinformation can cause models to fail at critical tasks—while still passing all standard benchmarks. Using the terrifying biology of prions as a guide, they investigate how adversarial attacks work, why they're so hard to detect, and what it means when we weaponize the vulnerabilities already baked into AI systems. ⏰ Key timestamps: 00:00 - The doctor's impossible choice: trusting an AI that lied 02:45 - The 2024 Nature Medicine study: 0.001% poisoned data, 99th percentile failure 05:15 - Connecting the dots: from Poison Gold (Ep 195) to adversarial ML 08:00 - Prions explained: how corrupted proteins cascade through biology 10:30 - The infection parallel: how poisoned training data spreads through models 13:00 - Why standard benchmarks miss the poison 15:15 - The unseen threat in AI we trust --- Sources & further reading: • Ian Goodfellow, Jonathon Shlens, Christian Szegedy — "Explaining and Harnessing Adversarial Examples" (2014) — [: https://arxiv.org/abs/1412.6572](https://arxiv.org/abs/1412.6572 • Daniel Junaid et al. — "Medical large language models are vulnerable to data-poisoning attacks" — Nature Medicine, 2024 — [: https://www.nature.com/articles/s41591-024-03445-1](https://www.nature.com/articles/s41591-024-03445-1 • Anthropic — "Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Training" (2024) — [: https://www.anthropic.com/research/sleeper-agents-training-deceptive-llms-that-persist-through-safety-training](https://www.anthropic.com/research/sleeper-agents-training-deceptive-llms-that-persist-through-safety-training • Anthropic — "Simple probes can catch sleeper agents" (2025) — [: https://www.anthropic.com/research/probes-catch-sleeper-agents](https://www.anthropic.com/research/probes-catch-sleeper-agents • Jiawei Su, Danilo Vasconcellos Vargas, Kouichi Sakurai — "One pixel attack for fooling deep neural networks" (2019) — [: https://arxiv.org/abs/1710.08864](https://arxiv.org/abs/1710.08864 • Anish Athalye, Logan Engstrom, Andrew Ilyas, Kevin Kwok — "Synthesizing Robust Adversarial Examples" (2018, ICML) — [: https://www.csail.mit.edu/news/why-did-my-classifier-just-mistake-turtle-rifle](https://www.csail.mit.edu/news/why-did-my-classifier-just-mistake-turtle-rifle • MIT Technology Review — "The AI lab waging a guerrilla war over exploitative AI" (2024) — [: https://www.technologyreview.com/2024/11/13/1106837/ai-data-posioning-nightshade-glaze-art-university-of-chicago-exploitation/](https://www.technologyreview.com/2024/11/13/1106837/ai-data-posioning-nightshade-glaze-art-university-of-chicago-exploitation/ • MIT Technology Review — "This tool strips away anti-AI protections from digital art" (2025) — [: https://www.technologyreview.com/2025/07/10/1119937/tool-strips-away-anti-ai-protections-from-digital-art/](https://www.technologyreview.com/2025/07/10/1119937/tool-strips-away-anti-ai-protections-from-digital-art/ • IEEE Spectrum — "In 2016, Microsoft's Racist Chatbot Revealed the Dangers of Online Conversation" — [: https://spectrum.ieee.org/in-2016-microsofts-racist-chatbot-revealed-the-dangers-of-online-conversation](https://spectrum.ieee.org/in-2016-microsofts-racist-chatbot-revealed-the-dangers-of-online-conversation • Microsoft Official Blog — "Learning from Tay's introduction" (2016) — [: https://blogs.microsoft.com/blog/2016/03/25/learning-tays-introduction/](https://blogs.microsoft.com/blog/2016/03/25/learning-tays-introduction/ • Andrew Ilyas, Shibani Santurkar, Dimitris Tsipras, Logan Engstrom, Brandon Tran, Aleksander Madry — "Adversarial Examples Are Not Bugs, They Are Features" (2019) — (NeurIPS 2019 proceedings) • …and 2 more in the episode research notes This podcast episode was fully generated by AI — research, script, voices, and production. Built with Claude, Piper TTS, and a
Listen Now

Love PodBriefly?

If you like Podbriefly.com, please consider donating to support the ongoing development.

Support Us