Podcast Episodes
Back to Search“Why study proto-training gaming as an adversarial alignment failure mode?” by Puria, Edward James Young, Cam
This is a dual post that lays out our current research project where we compare different pre-RL alignment methods and their ability to prevent mode…
3 weeks, 3 days ago
“Why study alignment interventions on pre-RL checkpoints?” by Edward James Young, Puria, Cam
This is a dual post that lays out our current research project where we compare pre-RL-training methods on their ability to prevent models from ‘pro…
3 weeks, 3 days ago
“Notes on technical alignment via human-like social drives” by Steven Byrnes
1. Frontmatter
1.1 Backstory for this post
As my regular readers know (see Intro to Brain-Like-AGI Safety), I’m working on the technical alignment p…
3 weeks, 3 days ago
“Reframing LessWrong-style decision theory as “commitment theory”” by Elias Schmied
Thanks to @eigengender and especially Chris Lakin and Simon Dima for valuable comments on a draft.
In this post, I will present an alternate framing…
3 weeks, 3 days ago
“The mosquito bucket of doom works” by dominicq
The mosquito bucket of doom is a population control mechanism where you dissolve some Bti (Bacillus thuringiensis israelensis) into a bucket and all…
3 weeks, 3 days ago
“Personascope: Measuring how deeply LLMs adopt personas” by Benji Berczi, Kyuhee Kim, Sid Black, Cozmin Ududec
Benji Berczi, Kyuhee Kim, James Requeima, Sid Black, Cozmin Ududec
This is work done by Benji and Kyuhee during MATS Winter 2026, mentored by Cozmin…
3 weeks, 3 days ago
“AI Safety Can’t Afford a Second Cause” by atlasaligned
Imagine an astronomer who discovers an asteroid with a 50% chance of hitting Earth in 2035. She goes on TV. She testifies before Congress. She found…
3 weeks, 3 days ago
“Superhuman Articulacy as an LLM Safety Target” by Dylan Bowman
TL;DR: Current LLMs are bad communicators relative to their agentic capabilities. I claim that articulacy is useful (and perhaps necessary) for AI s…
3 weeks, 4 days ago
“Experiments With Fabel’s Fiction” by Tomás B.
In keeping with my tradition and given I will lose access to Fabel in a couple days, I have asked fable to read all my organically written short sto…
3 weeks, 4 days ago
“No Space Like J-Space” by Zvi
There is a new very cool Anthropic paper: Verbalizable Representations Form a Global Workspace in Language Models. You can read the blog post veriso…
3 weeks, 4 days ago