Podcast Episodes
Back to Search"The Invisible Side of AI Governance" by Charbel-Raphaël
Tldr: Most strategic writing on AI governance on LessWrong describes the outsider game, which is most often visible: press, statements, open letters…
1 month ago
"A Theory of Prompt Injection (and why you should study roles)" by Charles Ye, softboiledheart
Summary
We've been building a theory of how prompt injections work under the hood.We show it comes down to how LLMs perceive roles (the humble chat …
1 month ago
"Machinic Psychopharmacology: Do LLMs Self-Medicate?" by Sid Black, Joseph Bloom
Sid Black, Joseph Bloom
UK AISI, Model Transparency Team
Epistemic status: Most experiments were run over a period of ~2-3 days during a hackathon a…
1 month ago
"Can activation verbalizers surface an internal chain of thought?" by oakhu, ryan_greenblatt
We introduce an evaluation for activation verbalizers: can they surface a target model's reasoning as it solves a math problem in a single forward p…
1 month ago
"The LLM shoggoth meme is weirder than you think" by HedonicEscalator
This article contains spoilers for At the Mountains of Madness, The Case of Charles Dexter Ward, and other works by H. P. Lovecraft.
In 1931, Claude…
1 month ago
[Linkpost] "Guardian Angels: LLM Personalization for Productivity and Security" by gwern
This is a link post. Powerful LLMs will be deployed at global scale in the next few years, and will dominate the Internet, and increasingly, ordinary…
1 month ago
"Gears for political races" by Tom Smith
In the past few years, many people around me have tried to convince me that US electoral politics is important. But like many other people in the co…
1 month ago
"A frontier AI company should shut down" by MichaelDickens
Cross-posted from my website.
Prior discussion: niplav's shortform (2025); Planning for Extreme AI Risks (2025) by Joshua Clymer
A frontier AI com…
1 month, 1 week ago
"Sympathy for both sides of the egregious misalignment debate" by Steven Byrnes
On one side of this debate is Yudkowsky & Soares, who think that (if AI progress continues) we’re on a direct path to egregiously-misaligned, schemi…
1 month, 1 week ago
"PSA: Almost nobody is working on alignment" by Chi Nguyen, peterbarnett
People often assume that a large fraction of the AI safety community works on alignment. As far as we're aware, this is not true. Most people are no…
1 month, 1 week ago