Podcast Episodes
Back to Search“How good are slop-vestigators?” by Hasan Baig, OscarGilg, Hamzah
TLDR:
We release MessageBoardAuditBench: a benchmark to measure how well agents can replicate the recent investigation into a swarm of OpenAI agents…1 week, 4 days ago
“Training on probes: What’s going on” by Charlie Steiner
TL;DR
If you train a probe for some property (like "honesty") and do gradient descent against this probe while continuing training that incentivizes…1 week, 4 days ago
“OpenAI have solved the Navier-Stokes Problem with a substantially more powerful model than Astra.” by fluxxrider
This is a linkpost for https://openai.com/index/navier-stokes-solution/
---
First published:
September 8th, 2026
1 week, 4 days ago
[Linkpost] “Frontier models still hack on simple variations of alignment evals from early 2025” by Dean Valentine
This is a link post.
In February 2025, back when o3-mini was the strongest available LLM, Palisade Research publicized a now well-known alignment eva…
1 week, 4 days ago
“Psychological Support for AI Safety Researchers Is Neglected and Easy to Provide” by Ihor Kendiukhov
I think there is a big chunk of relatively low-hanging-fruit-style neglected work useful for AI safety which I can roughly label as “psychological h…
1 week, 4 days ago
“Astra Is Hard to Monitor” by Zvi
OpenAI's central message on Astra is that it is three things:
Highly capable and can do all the things for you. Hard to monitor. The most align…1 week, 4 days ago
“Contra Piper on When Conversation Is Possible” by Zack_M_Davis
Kelsey Piper replies to Richard Ngo on Twitter:
since you have adopted the frame that liberals are self-deceived (and therefore not trustworthy abo…
1 week, 4 days ago
“An Alien Mind: Jakub Pachocki Warns Us” by Zvi
OpenAI Chief Scientist Jakub Pachocki is dropping truth bombs.
Tomorrow I will discuss Astra's lack of monitorability, and the potential contributi…
1 week, 5 days ago
[Linkpost] “Where are the token-level LLM kill-switches?” by beyarkay (Boyd Kane)
This is a link post.
Poisoned
Here's a simple idea: what if we trained in a string of characters that caused an LLM to emit the end of sequence token…
1 week, 5 days ago
“Machine Organizations” by Vaniver
OpenAI is nothing without its people
On November 20th, 2023, this was tweeted by many OpenAI employees as a sign of solidarity with Sam Altman in hi…
1 week, 5 days ago