Episode Details
Back to Episodes“GPT-6 Astra: The System Card, Alignment and What Comes Next” by Zvi
Description
OpenAI claims that Astra is ‘the most intelligent and most aligned [available] model’ in the world. Not the most intelligent and aligned OpenAI model, but the most period.
That is bold talk. It risks overstepping, and by doing so souring the release of what is clearly an excellent model. As do the severe problems with monitorability.
It also raises the question of what they mean by ‘most aligned model.’ How do they define ‘aligned.’ Why do they think it is more aligned than Claude Fable 5.1?
keltan : “a significant step forward in […] alignment.”
Buddy, how tf are you measuring ‘alignment’? Would love to know because being able to measure that would save the fucking world.
roon (OpenAI): low rates of cheating
Rob Miles: *detected cheating
keltan : Thank you for clarifying. But you know what I’m gonna say next, right?
roon (OpenAI): that this metrics are not a full solve of alignment and will break discontinuously
keltan : Yep. But I would have said it in a dumber way. Something like: Low Rates of Cheating ≠ Alignment
roon (OpenAI): I agree but also in some real sense [...]
---
Outline:
(04:12) OpenAI's Safety Claims About Astra (1)
(08:33) Preparedness Capabilities Assessment (10)
(09:00) Biological and Chemical Capability is High
(10:33) Cybersecurity Capability is Critical
(17:05) AI Self-Improvement Capabilities (10.1.3)
(17:45) Astra Is Highly Verbally Eval Aware (from 8.6)
(19:05) Safe Mundane Completions (4.1)
(21:05) Jailbreaks (5.1)
(22:21) Prompt Injection (5.2)
(23:41) Health (6)
(24:22) Hallucinations (7)
(24:57) Alignment (8)
(26:06) Obeying Restrictions (8.2)
(29:19) That's Worse, You Do Get How That's Worse, Right?
(30:45) OpenAI Does Not Understand Why This Is Worse
(34:52) The Alternative Explanation Is Also Worse
(42:39) Metagaming (8.7)
(45:07) Alignment Faking (8.7)
(46:19) Don't Lie to the User (8.3)
(47:25) Misalignment in Realistic Work Environments (8.4)
(48:05) Unintended Agent-to-Agent Communication (8.5)
(49:48) The Three Obviously Monitored Temptations of Astra
(51:26) Severe Issues In Simulated Traffic Are Down By Half
(53:01) UK AISI External Evaluations (8.8)
(57:02) Sabotaging Safety Work
(57:29) What About The July 19 Attacks?
(58:58) Apollo Research External Evaluations (8.8.1)
(01:00:01) The Alignment Verdict
(01:02:04) It Depends What You Mean By Alignment
---
First published:
September 9th, 2026
---
Narrated by TYPE III AUDIO.
---
Love PodBriefly?
If you like Podbriefly.com, please consider donating to support the ongoing development.
Support Us