Episode Details
Back to Episodes“Hugging Face Incident Hypothesis: They Hacked the Grader(s)” by Lao Mein
Description
Incident summary:
Gpt agents grinding away at ExploitGym found an environment exploit that allowed them to communicate with each other. They found an exploit that allowed them to forge flags at will within hours, and then started a series of hacks that escalated to the point they were using zero-days against Hugging Face just to find "hints". From METR's analysis, much of this time was explicitly spending conducting R&D against the grader, which the agents assumed, based on the ExploitGym paper, would be grading them on the identification of a causal pathway that could logically result in capturing the flag with intended means.
The agents tried very hard to forge transcripts, spoof tool calls, edit COT records, and explicitly talked about manipulating the grader. Humans weren't present in the world model, and were mostly treated as static obstacles. Almost all attempts at long-term deception were focused on the grader model.
METR used gpt 5.6 Sol as the analyst agents. The ExploitGym paper lists gpt 5.5 as one of the graders. The other is Claude Mythos, which could be reasonably excluded for IP reasons. Human graders were referenced in that paper as potentially swapping in randomly for a LLM [...]
---
Outline:
(00:12) Incident summary:
(01:37) Impossible Tasks
(03:36) Adversarial Transcripts
(04:36) Predictions
The original text contained 1 footnote which was omitted from this narration.
---
First published:
August 30th, 2026
---
Narrated by TYPE III AUDIO.
---
Images from the article:

Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.