Podcast Episodes
Back to Search“Announcing the Safe Pareto Improvements (SPI) Fundamentals Program” by Anthony DiGiovanni
CLR is excited about safe Pareto improvements (SPIs) as a way to mitigate downsides from conflict between AIs. SPIs are a class of interventions on …
4 weeks, 1 day ago
“You Should Choose How You React to Your Feelings” by Nate Sharpe
One of the things I love about parenting is being frequently reminded how many things are not the default for most people. While I see this most cle…
4 weeks, 1 day ago
“Lydia Laurenson: “The Inside Story of Leverage Research”” by Davis_Kingsley
Lydia Laurenson recently posted an article called "The Inside Story of Leverage Research" that gets into substantially more detail on what went on i…
4 weeks, 1 day ago
“When Role-playing, Do Models Believe What They Say?” by Sturb, David Africa, Sid Black
TL;DR
When a model role-plays a persona, does it only change what it says, or also what it internally represents as true?To study this, we induce pe…4 weeks, 1 day ago
“I can’t think of good interventions for ensuring third-party model access.” by Cleo Nardo
Summary
I'm increasingly convinced that model access parity is a big deal and we are not on track to achieve it. By model access parity, I mean a sm…
4 weeks, 1 day ago
“Research update: RL on Debate Games shows Proposal Accuracy uplift alongside Judge Hacking” by lennie, joanv, Shi, Jacob Pfau
The first three sections are written for a general TAIS reader who wants to understand what the state of Debate research is and some high-level take…
4 weeks, 1 day ago
“AI #175: The Fable Continues” by Zvi
Fable's back. Back again. Fable's back. Tell a friend. Use your free week to its fullest.
This is excellent news. The blip only lasted a few weeks.…
4 weeks, 1 day ago
“Conversation Among Cade Metz, Michael Vassar, Jessica Taylor, and Zack M. Davis” by Zack_M_Davis
(Previously, previously, previously.)
20–21 August 2025
From: Zack M. Davis
To: Cade Metz
CC: Benjamin Hoffman, Jessica Taylor, Michael Vassar
…
4 weeks, 2 days ago
“AI Futurism Reading List” by Alexa Pan
We at Redwood recently ran a strategy fellowship through Astra. As part of this, we ran a reading group for our fellows on some of the topics that w…
4 weeks, 2 days ago
“AI welfare research needs basic science” by OscarGilg, Pierre Beckmann, Jake1638
Over the course of MATS 9.0 we formed some views about AI welfare research that we thought were worth writing up. This post is meant to spark discus…
4 weeks, 2 days ago