Episode Details

Back to Episodes

“11 Open Empirical Problems in Reward-Seeking” by Alex Meinke, Jérémy Scheurer, Axel Højmark, Theodore Ehrenborg

Published 1 week, 3 days ago
Description

We recently published our paper on "Measuring Reward-Seeking via Contrastive Belief Updates". We're excited about research like this, and there are many more open problems than we can work on. Here's a list of open problems that we think are valuable.

If you work on/solve these problems, we'd be happy to signal-boost your research. If your next research project is one of these problems, feel free to reach out to alex@apolloresearch.ai to discuss it in more detail.

Reward-Seeking and its Implications

1. Is a Reward-Seeking Model more difficult to align?

The strongest case for expecting reduced "train-time corrigibility" due to reward-seeking, applies to Instrumental Reward-Seeking, where the model actively reasons "I will please oversight now, in order to accomplish some other thing later". Alignment training a model like that may update its beliefs about graders and oversight, without reshaping its underlying values.

Current forms of reward-seeking are likely better understood as terminal, i.e. models try to please the grader without ulterior motives. There is likely a continuous spectrum between Terminal and Instrumental Reward-Seeking. Thus, we can hopefully study the effects that mostly Terminal Reward-Seeking has on train-time corrigibility now, in the hopes of learning about the effects that Instrumental [...]

---

Outline:

(00:43) Reward-Seeking and its Implications

(00:47) 1. Is a Reward-Seeking Model more difficult to align?

(02:20) 2. Can we measure Instrumental Reward-Seeking?

(03:25) 3. When does Reward-Seeking most increase / decrease?

(04:50) 4. Are there better Belief Update Techniques than SDF?

(07:06) Improving SDF

(07:10) 5. Can AIs detect the difference between pretraining facts and SDF facts?

(07:49) 6. How is SDF different from changing pretraining?

(09:32) 7. Does SDF have off-target effects?

(11:29) 8. SDF sometimes generalizes in surprising ways

(12:31) 9. Getting good synthetic document recall is finicky

(14:11) Better Validation Techniques

(14:23) 10. Our model organisms could be more robust

(15:32) 11. Can we get more bits of information for ground-truth?

The original text contained 1 footnote which was omitted from this narration.

---

First published:
July 21st, 2026

Source:
https://www.lesswrong.com/posts/8wXRuHQqCbRsbap6q/11-open-empirical-problems-in-reward-seeking

---

Narrated by TYPE III AUDIO.

---

Images from the article:

Diagram showing

Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

Listen Now

Love PodBriefly?

If you like Podbriefly.com, please consider donating to support the ongoing development.

Support Us