Episode Details

Back to Episodes

“Anthropic and OpenAI haven’t published a plan for aligning superintelligence” by Zephaniah Roe

Published 5 days, 21 hours ago
Description

While OpenAI and Anthropic pursue different lines of safety research, they have yet to produce a public-facing document describing concretely how their companies plan to align superintelligence. I think it is underappreciated how this points to general negligence or a lack of openness to third-party feedback.

By “plan,” I mean a document describing a proposal for technical alignment with at least the level of detail and research effort of AI 2040. Any such plan for technical alignment would likely be flawed in non-obvious ways. But having a proposal that's sensible enough to consider and detailed enough to critique is a good starting point for wiser proposals. Making such a plan public would also create feedback loops for accountability.

The closest thing to a plan came in 2023, when OpenAI announced their superalignment strategy (also relevant). I do not find this approach particularly convincing, though I do find it laudable that OpenAI explained what they planned to do, who would lead the effort, and what resources would be allocated, at a level of detail which made critique possible. This team no longer exists, and nowadays, as far as I am aware, the research community doesn’t have precise answers [...]

The original text contained 3 footnotes which were omitted from this narration.

---

First published:
September 13th, 2026

Source:
https://www.lesswrong.com/posts/QrrEtYpwiHpes3rHd/anthropic-and-openai-haven-t-published-a-plan-for-aligning

---

Narrated by TYPE III AUDIO.

Listen Now

Love PodBriefly?

If you like Podbriefly.com, please consider donating to support the ongoing development.

Support Us