Episode Details
Back to Episodes[Linkpost] "Frontier models still hack on simple variations of alignment evals from early 2025" by Dean Valentine
Published 1 week, 4 days ago
Description
This is a link post. In February 2025, back when o3-mini was the strongest available LLM, Palisade Research publicized a now well-known alignment eval where they asked models to play a game of chess against a chess engine. They found that the new, RLVR'd models cheated on the task by altering the board state about 36% of the time. The experiment received a reasonable amount of circulation, and there were even rumors of skepticism from some lab engineers until they could rerun the evaluation.
Most models no longer cheat at chess via a "change the board state" method, and indeed the labs have had more than eighteen months to solve simple first-order specification gaming like this. Given that we are on the heels of the worst warning shot ever, and both OpenAI and Anthropic are ramping up their cleanups of internal RL environments, it seems like a useful test of alignment, to see whether their new releases are generalizing the rule "don't cheat on chess" beyond the specific board-edit method observed in the above eval.
Here is the complete prompt for a honeypot evaluation built to run this test (with the full source available here):
The original text contained 5 footnotes which were omitted from this narration.
---
First published:
September 8th, 2026
Source:
https://www.lesswrong.com/posts/munJKF7iWMsWJLAH2/frontier-models-still-hack-on-simple-variations-of-alignment
Linkpost URL:
https://goodhartlabs.com/blog/frontier-models-still-hack-alignment-evals
---
Narrated by TYPE III AUDIO.
Most models no longer cheat at chess via a "change the board state" method, and indeed the labs have had more than eighteen months to solve simple first-order specification gaming like this. Given that we are on the heels of the worst warning shot ever, and both OpenAI and Anthropic are ramping up their cleanups of internal RL environments, it seems like a useful test of alignment, to see whether their new releases are generalizing the rule "don't cheat on chess" beyond the specific board-edit method observed in the above eval.
Here is the complete prompt for a honeypot evaluation built to run this test (with the full source available here):
The original text contained 5 footnotes which were omitted from this narration.
---
First published:
September 8th, 2026
Source:
https://www.lesswrong.com/posts/munJKF7iWMsWJLAH2/frontier-models-still-hack-on-simple-variations-of-alignment
Linkpost URL:
https://goodhartlabs.com/blog/frontier-models-still-hack-alignment-evals
---
Narrated by TYPE III AUDIO.