Episode Details

Back to Episodes

“Reward Laundering: LLMs Can Gain Unintended Behaviors by Deciding When to Earn Their Rewards” by egan, abhayesian, Jozdien

Published 13 hours ago
Description

This work was done by an automated research scaffold developed at Redwood Research. abhayesian provided the initial project idea. The agent designed and ran all experiments and produced a detailed writeup, which humans (with AI assistance) distilled into this more readable post.

We think this project is at the level of rigor of a mid-MATS research update. We assessed the correctness mostly by looking at the writeups to make sure that things like the experiment design making sense baselines being reasonable. We didn't do detailed code reviews, aside from running an automated LLM reviewer and spot checking that the final codebase's results were consistent, but we release the codebase. More details about LLM usage in the Appendix.

TL;DR – We find that Qwen 3.5 9B can utilize its RL training process on one task to self-improve at another. By choosing to earn reward on the easy, trained task only when it also performs the hard task well, Qwen can train itself on a hard, easily verifiable task that is never directly rewarded.

💻 Codebase

Introduction

Exploration hacking refers to a set of threat models where a model strategically alters its exploration during RL training in order to influence the [...]

---

Outline:

(01:23) Introduction

(02:52) Setup

(04:57) Results

(07:32) Discussion

(09:45) Appendix

(09:48) AI Involvement With The Project

(12:21) Prompts

(12:34) Example rollouts

The original text contained 3 footnotes which were omitted from this narration.

---

First published:
July 31st, 2026

Source:
https://www.lesswrong.com/posts/fPWP4rHPLqKKHKe6B/reward-laundering-llms-can-gain-unintended-behaviors-by

---

Narrated by TYPE III AUDIO.

---

Images from the article:

Scatter plot titled
Line graphs comparing

Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

Listen Now

Love PodBriefly?

If you like Podbriefly.com, please consider donating to support the ongoing development.

Support Us