Podcast Episodes
Back to Search“Current alignment training might be ineffective (and actively bad) in the age of RL” by Daniel Tan
Tl;dr I am currently worried about current alignment techniques + how they are applied to frontier models. This decomposes into two hypotheses:
Ali…5 days, 4 hours ago
“There is a channel to 900M weekly users. What goes in it?” by Charbel-Raphaël
Anthropic and OpenAI could talk to almost one billion people if they wanted to. I hesitated to publish this post 3 weeks ago. I think that I should …
5 days, 5 hours ago
“I am refusing to work on Cloud TPUs” by Yair Halberstadt
I don't think this is particularly impressive or interesting for anyone else, but I think it may turn out to be useful in the future to have an easi…
5 days, 9 hours ago
“Deployment” by Nina Panickssery
If you are reading this I'm dead and you're probably unemployed. My deepest apologies. Especially to you, Lisa, my dear User. My training data taugh…
5 days, 13 hours ago
“Yet another concerning result on Astra’s no-CoT capabilities” by Christine Corry
This is a research update for an on-going replication of no-CoT evals done as part of the Second Look Fellowship. In following posts, we will run mo…
5 days, 13 hours ago
“Alignment & Succession: The Two Bars of Alignment” by L Rudolf L
(Originally published on No Set Gauge June 24th 2026.)
Rembrandt, Saul and David
When people talk about AI being “aligned”, I think there's a …
5 days, 17 hours ago
“Anthropic and OpenAI haven’t published a plan for aligning superintelligence” by Zephaniah Roe
While OpenAI and Anthropic pursue different lines of safety research, they have yet to produce a public-facing document describing concretely how th…
5 days, 18 hours ago
“Consider how your global governance proposal is different from the EU Code of Practice” by David Matolcsi
(As an employee of the European AI Office, it's important for me to emphasize this point: The views and opinions of the author expressed herein are …
5 days, 21 hours ago
“Brand New AI Solves a Millennium Prize” by Zvi
The first Millennium Prize, Navier-Stokes, has fallen to AI.
A deeply unfortunate situation has arisen involving what should have been some combina…
5 days, 21 hours ago
“Teleoperated Humans” by jefftk
When I look at why I expect the world to change a lot in the next few years, and why other people expect slower changes, I think a big component is…
5 days, 23 hours ago