Podcast Episodes

Back to Search
“How I’m Evaluating Corrigibility Grant Applications” by Max Harms

I'm the sole manager of the newly created Corrigibility Research Fund. While I've been an alignment researcher for a long time, this is my first tim…

2 weeks, 3 days ago

Short Long
View Episode
“From safety research prompt to cross-model universal jailbreak” by richbc

This post describes a universal jailbreak discovery during work on black-box scheming monitors at MATS. The jailbreak itself is not released; see On…

2 weeks, 3 days ago

Short Long
View Episode
“Cat-Belling Problems” by Eliezer Yudkowsky

(Originally written in 2021, if the discussion around AI now seems odd; it is written for a time when people were still trying to solve what would n…

2 weeks, 3 days ago

Short Long
View Episode
“Steering towards “automated grading” degrades alignment” by Jan Betley, Johannes Treutlein, Clément Dumas

TL;DR: We steer Qwen3.6-27B on a dimension constructed from the contrast pair “a script will verify your answer” (automated grader) vs “a human will…

2 weeks, 3 days ago

Short Long
View Episode
“Invididual Effort to Reduce Biorisk” by jefftk

I'm pretty worried about how AI might change the world a lot very soon. In some of these cases things go very wrong very quickly, in others things…

2 weeks, 3 days ago

Short Long
View Episode
[Linkpost] “Sen. Bernie Sanders (I-VT) and Rep. Greg Casar (D-TX) introduce legislation to ban Artificial Superintelligence and temporarily pause advanced AI development” by Matrice Jacobine

This is a link post.

[...]

“Nearly every day, there is a frightening new story about how Big Tech companies are losing control of the technology they…

2 weeks, 3 days ago

Short Long
View Episode
“What is neuralese and why is it bad?” by Linch

What is neuralese?

To explain neuralese, we need to first understand chain-of-thought, one of the largest developments in AI in the last five years.…

2 weeks, 3 days ago

Short Long
View Episode
“A proposal for a highly effective AI safety org” by ceselder

TLDR: an org that pays people to "just read the fucking transcripts"; a large amount of people reading anonymized claude code/RL/eval transcripts fl…

2 weeks, 4 days ago

Short Long
View Episode
“Talking to journalists” by KatjaGrace

A common view around me seems to be that journalists are frequently dishonorable and dangerous, and talking to them is a risk to be avoided unless y…

2 weeks, 4 days ago

Short Long
View Episode
[Linkpost] “Resolution has a new Agent Foundations team” by Jeremy Gillen

This is a link post.

The team will include me (Jeremy Gillen), Abram Demski, Sam Eisenstat, Scott Garrabrant and Kaarel Hänni. We'll soon recruit add…

2 weeks, 4 days ago

Short Long
View Episode

Love PodBriefly?

If you like Podbriefly.com, please consider donating to support the ongoing development.

Support Us