Episode Details
Back to Episodes“AGI Safety and Alignment at Google DeepMind: A Summary of Recent Work (July 2026)” by Rohin Shah, Seb Farquhar
Description
It's been nearly two years since our last major update here in August 2024 and we wanted to share another recap of our recent work with the AGI safety community. Things have changed a lot since then. We are now fully in the midgame, and focus more on landing things in production.
Who are we?
We are the AGI Safety and Alignment Team (ASAT), the main group at Google DeepMind working directly on technical approaches to existential risk from AI systems. Last year we published An Approach to Technical AGI Safety and Security, which remains the best place to read our overarching vision.
Highlights
Norms around chain of thought. Our impression is that our work meaningfully moved the field away from beliefs along the lines of “chain of thought is often unfaithful and so not worth using” towards beliefs along the lines of “chain of thought is a very useful tool that is worth preserving”, leading to a tentative industry consensus on its importance. We have also published substantial technical research that enables companies to preserve chain of thought transparency for longer than would have happened by default.
We think this is a big deal: extending the period [...]
---
Outline:
(00:33) Who are we?
(00:54) Highlights
(02:43) Agent Control & Monitorability
(05:47) Deep Alignment
(07:31) Language Model Interpretability
(09:59) Amplified Oversight
(11:57) Alignment Evaluations
(13:35) Frontier Safety: Risk Assessment & Mitigations
(16:29) Causal Alignment
(17:14) External advising
---
First published:
July 31st, 2026
---
Narrated by TYPE III AUDIO.