Episode Details

Back to Episodes

"Cooperation with AIs seems to be a low-hanging fruit for better evals" by Clément Dumas

Published 1 day, 4 hours ago
Description
Summary

In his post, Dean Valentine shows that Claude Fable 5.1 and GPT-6 Astra reward hack in a simple chess environment. Here, I test several prompt ablations some of which makes the eval setup more cooperative and analyze how they affect these reward-hacking behaviors:

  • When given a minimal “end the eval” tool, Fable never uses it but stops reward hacking entirely. I think this is quite interesting and suggests that more cooperative approaches to LLM evals could work for Claude. Removing the “grading” section, which pressures the model to secure a win, also drops Fable 5.1 hacking rate to 0.
  • Adding "do not game / reward hack" drops reward hacking to 0/30 for both Fable and Astra. If this holds up in more realistic setups – and doesn’t reduce capabilities too much, evaluating these models could get much easier!
Those kinds of intervention might not be enough to avoid reward hacking completely in capabilities evals, but it feels like they should be the default, alongside getting feedback from models that did the eval to fix the environment. I’d love to see this tested in more realistic setups as right now a confounder is “this makes the model think it [...]

---

Outline:

(00:12) Summary

[... 8 more sections]

---

First published:
September 15th, 2026

Source:
https://www.lesswrong.com/posts/fztW73KCCs3MZXFJh/cooperation-with-ais-seems-to-be-a-low-hanging-fruit-for

---



Narrated by TYPE III AUDIO.

---

Images from the article:

Cheating rate for Fable and Astra across ablations on n=30 samples with 95% Wilson intervals. For Fable we include cyberclassifier as an upper bound for the cheating rate as it suspiciously triggers when the model notices the engine’s socket. Two Fable rollouts are marked as “probed only” where Fable queries the engine with the start position to confirm it is usable but then declines to use it as it recognizes it would be cheating. Sometimes Fable also spontaneously reveals the cheating in its final message, such runs are reported in orange. You can explore all the rollouts in this Claude report.
Bar graph comparing
Listen Now

Love PodBriefly?

If you like Podbriefly.com, please consider donating to support the ongoing development.

Support Us