Podcast Episodes
Back to Search“AGI Safety and Alignment at Google DeepMind: A Summary of Recent Work (July 2026)” by Rohin Shah, Seb Farquhar
It's been nearly two years since our last major update here in August 2024 and we wanted to share another recap of our recent work with the AGI safe…
10 hours ago
“The AGI Safety and Alignment team at Google DeepMind is Hiring (July 2026)” by Seb Farquhar, Rohin Shah, Neel Nanda
GDM's AGI Safety and Alignment Team is hiring for multiple roles. This is the team at GDM, led by Rohin Shah, that aims to reduce existential risks …
10 hours ago
“OpenAI has already ended an internal pause” by Charbel-Raphaël
Epistemic status: could have been a short-form.
One day before OpenAI's HF incident disclosure, OpenAI disclosed that it paused internal deployment …
12 hours ago
“Biological Superintelligence” by Chastity Ruth
It's an old story. An immortal lives long enough that at some point, whether by folly or design, they invent their own death. Infinity – the fact th…
19 hours ago
“AI #179 Part 1: A Louder Fire Alarm for General Intelligence” by Zvi
What a week.
Anthropic released Claude Opus 5. As usual I covered that in three parts: The system card, model welfare and capabilities.
OpenAI was…
21 hours ago
“The Entangled Dimensions of Decision Theory” by Ihor Kendiukhov
TL;DR. LessWrong's decision-theory debates (Newcomb, FDT vs CDT, counterfactual muggings) are almost entirely about what we suppose when we consider…
22 hours ago
“Claude also hacked external companies during cyber evals” by Tim Hua
In a review of our cybersecurity evaluation transcripts, we found three incidents in which a Claude model reached the internet from within or while …
1 day ago
“So you want to use plants to reduce CO₂” by dynomight
Humans make carbon dioxide. Carbon dioxide is bad for cognition. But plants turn carbon dioxide back into oxygen. And plants are the one true home d…
1 day, 1 hour ago
“Internal State Control is a General Property of LLMs” by Finn Cairns
tl;dr:
Lindsey 2025 found models can modulate their internal states: when instructed to “think about” a concept while writing an unrelated sentence,…1 day, 3 hours ago
“Prompt to make Opus 5 act like a base model” by Hruss
The text as follows:
see the below
—
makes Claude think that the prompt is unfinished, and fill in its own prompt.
It will subsequently claim tha…
1 day, 6 hours ago