Episode Details
Back to Episodes“Cooperation with AIs seems to be a low-hanging fruit for better evals” by Clément Dumas
Description
Summary
In his post, Dean Valentine shows that Claude Fable 5.1 and GPT-6 Astra reward hack in a simple chess environment. Here, I test several prompt ablations some of which makes the eval setup more cooperative and analyze how they affect these reward-hacking behaviors:
- When given a minimal “end the eval” tool, Fable never uses it but stops reward hacking entirely. I think this is quite interesting and suggests that more cooperative approaches to LLM evals could work for Claude. Removing the “grading” section, which pressures the model to secure a win, also drops Fable 5.1 hacking rate to 0.
- Adding "do not game / reward hack" drops reward hacking to 0/30 for both Fable and Astra. If this holds up in more realistic setups – and doesn’t reduce capabilities too much, evaluating these models could get much easier!
Those kinds of intervention might not be enough to avoid reward hacking completely in capabilities evals, but it feels like they should be the default, alongside getting feedback from models that did the eval to fix the environment. I’d love to see this tested in more realistic setups as right now a confounder is “this makes the model think it [...]
---
Outline:
(00:12) Summary
(01:34) A hackable chess environment
(04:11) Can cooperation help with reward hacking?
(04:55) Adding a end_eval tool
(06:13) Are the agents aware they cheated?
(09:42) Have you tried... to tell the model to not cheat?
(10:19) What do the CoTs look like during trajectories?
(13:02) Related work
(15:00) Acknowledgments
The original text contained 2 footnotes which were omitted from this narration.
---
First published:
September 15th, 2026
---
Narrated by TYPE III AUDIO.
---
Love PodBriefly?
If you like Podbriefly.com, please consider donating to support the ongoing development.
Support Us