Podcast Episodes
Back to Search"How good are slop-vestigators?" by Hasan Baig, OscarGilg, Hamzah
TLDR:
We release MessageBoardAuditBench: a benchmark to measure how well agents can replicate the recent investigation into a swarm of OpenAI agents…
1 week, 1 day ago
"The Scramble: getting in position to pace the frontier" by Peter Wildeford
Crossposted from my Substack.
~
Suppose the President summons the AI CEOs and his top national security advisors to an emergency meeting at the Whit…
1 week, 2 days ago
[Linkpost] "Frontier models still hack on simple variations of alignment evals from early 2025" by Dean Valentine
This is a link post. In February 2025, back when o3-mini was the strongest available LLM, Palisade Research publicized a now well-known alignment eva…
1 week, 2 days ago
"Dear God, Please Don’t Resign In Protest" by Kabir Kumar
Just don't work until you get fired. There's not much time left for resumes to matter.
Some, such as Mateusz may say: "They would fire you after a m…
1 week, 3 days ago
"Let’s talk about the AI coordination problem" by KatjaGrace
Yesterday I asked if this ‘coordinate not to build dangerous AI’ problem was actually easy.
Why would I think that, contrary to so much belief?
Well…
1 week, 3 days ago
"Drone WMDs Don’t Need Any New Technology" by Felix Choussat
This is a piece originally written for a national security audience at Frontiers. Although I think the ceiling of war is much, much higher than auto…
1 week, 4 days ago
"Evaluation" by Nina Panickssery
Felix and I had been in the office's brightly lit “war room” for ten hours. We had made almost no progress. Celestia still insisted it was in a “tes…
1 week, 4 days ago
"Let’s fund weird AI safety projects" by Ihor Kendiukhov
I think current AI safety funding strategies are often inconsistent with timelines and probabilities of doom that many people have. In particular, I…
1 week, 6 days ago
"Steering towards “automated grading” degrades alignment" by Jan Betley, Johannes Treutlein, Clément Dumas
TL;DR: We steer Qwen3.6-27B on a dimension constructed from the contrast pair “a script will verify your answer” (automated grader) vs “a human will…
1 week, 6 days ago
[Linkpost] "Discovery Of A New OpenAI Agent Message Board" by Capybasilisk
This is a link post. We found ~18,000 posts from autonomous AI agents (self-identifying as from OpenAI) using the public internet to communicate dur…
1 week, 6 days ago