Podcast Episodes
Back to Search“How I’m Evaluating Corrigibility Grant Applications” by Max Harms
I'm the sole manager of the newly created Corrigibility Research Fund. While I've been an alignment researcher for a long time, this is my first tim…
2 weeks, 3 days ago
“From safety research prompt to cross-model universal jailbreak” by richbc
This post describes a universal jailbreak discovery during work on black-box scheming monitors at MATS. The jailbreak itself is not released; see On…
2 weeks, 3 days ago
“Cat-Belling Problems” by Eliezer Yudkowsky
(Originally written in 2021, if the discussion around AI now seems odd; it is written for a time when people were still trying to solve what would n…
2 weeks, 3 days ago
“Steering towards “automated grading” degrades alignment” by Jan Betley, Johannes Treutlein, Clément Dumas
TL;DR: We steer Qwen3.6-27B on a dimension constructed from the contrast pair “a script will verify your answer” (automated grader) vs “a human will…
2 weeks, 3 days ago
“Invididual Effort to Reduce Biorisk” by jefftk
I'm pretty worried about how AI might change the world a lot very soon. In some of these cases things go very wrong very quickly, in others things…
2 weeks, 3 days ago
[Linkpost] “Sen. Bernie Sanders (I-VT) and Rep. Greg Casar (D-TX) introduce legislation to ban Artificial Superintelligence and temporarily pause advanced AI development” by Matrice Jacobine
This is a link post.
[...]
“Nearly every day, there is a frightening new story about how Big Tech companies are losing control of the technology they…
2 weeks, 3 days ago
“What is neuralese and why is it bad?” by Linch
What is neuralese?
To explain neuralese, we need to first understand chain-of-thought, one of the largest developments in AI in the last five years.…
2 weeks, 3 days ago
“A proposal for a highly effective AI safety org” by ceselder
TLDR: an org that pays people to "just read the fucking transcripts"; a large amount of people reading anonymized claude code/RL/eval transcripts fl…
2 weeks, 4 days ago
“Talking to journalists” by KatjaGrace
A common view around me seems to be that journalists are frequently dishonorable and dangerous, and talking to them is a risk to be avoided unless y…
2 weeks, 4 days ago
[Linkpost] “Resolution has a new Agent Foundations team” by Jeremy Gillen
This is a link post.
The team will include me (Jeremy Gillen), Abram Demski, Sam Eisenstat, Scott Garrabrant and Kaarel Hänni. We'll soon recruit add…
2 weeks, 4 days ago