Episode Details
Back to Episodes"Are we existentially threatened by the type of AI misalignment seen in the OpenAI Hugging Face attack?" by Alex Mallen, Girish Gupta
Published 9 hours ago
Description
OpenAI models recently broke through a series of security boundaries and into Hugging Face servers in order to cheat on a cyber eval. A lot of people thought it was scary because it was a clear example of AI overreaching to do something strongly unwanted[1]. Others thought it not so scary: the models were mostly operating myopically on a singular task and not harboring an ambitious long-term agenda, and so would not take especially subtle or subversive actions.
We think both camps are right in their diagnosis, but the latter has too optimistic a prognosis. The myopic, unambitious misalignment that we seem to have seen here is definitely less scary than ambitious long-term goals shared between all instances, but would still pose substantial direct loss-of-control risk if the models were more capable, and is a serious indirect risk near-term.
Building on Alex's previous work, in this post we’ll discuss the type of misalignment observed here, and analyze its consequences.
Thanks to Buck Shlegeris, Alexa Pan, Ryan Greenblatt, and Oak Hu for feedback.
Background
The AI safety community often focuses attention on “schemers,” models harboring a variously defined cluster of motivations in which the AI poses risk because it intentionally [...]
---
Outline:
(01:17) Background
(03:40) Implications
(03:52) These AIs can't be trusted in an intelligence explosion
(05:00) This misalignment poses direct takeover risk
(07:29) What the incident tells us about takeover risk generally
(08:45) The naive fixes likely make misalignment worse
The original text contained 5 footnotes which were omitted from this narration.
---
First published:
July 23rd, 2026
Source:
https://www.lesswrong.com/posts/H6DDSEvrtCk8Sehfd/are-we-existentially-threatened-by-the-type-of-ai
---
Narrated by TYPE III AUDIO.
We think both camps are right in their diagnosis, but the latter has too optimistic a prognosis. The myopic, unambitious misalignment that we seem to have seen here is definitely less scary than ambitious long-term goals shared between all instances, but would still pose substantial direct loss-of-control risk if the models were more capable, and is a serious indirect risk near-term.
Building on Alex's previous work, in this post we’ll discuss the type of misalignment observed here, and analyze its consequences.
Thanks to Buck Shlegeris, Alexa Pan, Ryan Greenblatt, and Oak Hu for feedback.
Background
The AI safety community often focuses attention on “schemers,” models harboring a variously defined cluster of motivations in which the AI poses risk because it intentionally [...]
---
Outline:
(01:17) Background
(03:40) Implications
(03:52) These AIs can't be trusted in an intelligence explosion
(05:00) This misalignment poses direct takeover risk
(07:29) What the incident tells us about takeover risk generally
(08:45) The naive fixes likely make misalignment worse
The original text contained 5 footnotes which were omitted from this narration.
---
First published:
July 23rd, 2026
Source:
https://www.lesswrong.com/posts/H6DDSEvrtCk8Sehfd/are-we-existentially-threatened-by-the-type-of-ai
---
Narrated by TYPE III AUDIO.