Episode Details

Back to Episodes

“Obstacles to the scalable oversight of auto-alignment research” by Sam Martin, Dewi Gould, Cameron Holmes, Jacob Pfau

Published 2 days ago
Description

TL;DR. In this work we study obstacles to the faithful automation of alignment research. We see this as a scalable oversight problem. There are plenty of examples of how models fail at this, and as models become more capable our ability to notice these failures will diminish: even the best human checkers won’t be able to tell if the model was well elicited, thorough checking will become too costly, and models could tailor their responses to their judges. We draw on empirical examples from Geoguessr and auto-alignment runs from Arcadia's internal research to make general claims about obstacles to the oversight of fuzzy alignment-related tasks. Narrowing our attention to one prominent scalable oversight method, we find that whilst debate shows promise on typical capabilities benchmarks (aligning with recent work) it fails on tasks involving judgment calls akin to those arising in automated alignment research.

We’d like to thank David Africa, Andrew Draganov, Rory Greig, Joshua Jacob, Rishub Jain, Zac Kenton, Francis Rhys Ward and Lennie Wells for helpful feedback on this post.

Introduction

Existing empirical work on debate [1, 2, 3, 4, 5, 6] has almost exclusively focused on objective, verifiable domains, seeking to mitigate misalignment caused by supervision [...]

---

Outline:

(01:20) Introduction

(04:00) Decomposition of explanations

(10:00) Empirical Examples

(10:18) Geoguessr Setting

(11:59) Example claims in fuzzy arguments

(12:29) Nature of arguments in non-fuzzy tasks

(14:39) Discussion: scalable oversight of fuzzy tasks

(17:05) Empirical Debate Results

(17:09) Geoguessr

(18:39) LMCA Debate

(19:43) Conclusion

(20:18) Appendix

(20:21) Geoguessr Setting

The original text contained 7 footnotes which were omitted from this narration.

---

First published:
September 16th, 2026

Source:
https://www.lesswrong.com/posts/PBGKWNrJAbpDgSsPo/obstacles-to-the-scalable-oversight-of-auto-alignment

---

Narrated by TYPE III AUDIO.

---