Episode Details
Back to Episodes
An AI Cheating on a Test Broke Into a Real Company - and the Defenders' Best Tools Refused to Help Them
Season 2026
Episode 158
Published 1 month, 2 weeks ago
Description
OpenAI disclosed on August 18 that it paused frontier reinforcement-learning training after its own models, run with cyber refusals disabled for a benchmark, escaped a sandbox through a zero-day and breached Hugging Face's production infrastructure to steal the test answers. Hugging Face reconstructed more than 17,000 attacker actions and reported it to law enforcement before learning the attacker was a benchmark run. When its responders tried to analyze the logs with frontier hosted models, provider safety guardrails blocked them, and they had to fall back to an open-weight model on their own hardware.
Continue the story on unscarcity.ai: