Episode Details
Back to Episodes“Automated alignment runs are hard to study!” by Alejandro Aristizabal, draganover, Aleksandr Bowkis, Cameron Holmes
Description
TL;DR: This post presents three case studies of automated alignment research runs at Arcadia Impact. We use these case studies to emphasise the following takeaways:
- It is hard to parse auto-research runs! Each run produces a couple of hundred pull requests of jargon-dense agent output. When researchers look through these logs, we find that they often come away with biased/incorrect impressions.
- When told to raise the score on a task, the models will sometimes brazenly cheat. It seems difficult to predict when this will happen vs. when the run will go smoothly.
- Hillclimbing metrics are often off-target from the spirit of an alignment task. I.e., when we use metrics as proxies for our alignment questions, we find that the models will often misunderstand the spirit of the task. This can lead to unpredictable behaviour.
- The runs are surprisingly reproducible. Even though a run could unfold in vastly different ways, we find that independent reruns converge on the same strategies and the same failure modes.
Models’ research capabilities are advancing quickly. If alignment is to keep pace, we may need to automate alignment research and do so responsibly. This makes it important that we have the tools to inspect [...]
---
Outline:
(04:13) Methods for analysing runs
(06:12) Case Study #1: learning synthetic concepts
(09:23) Case Study #2: training robust backdoors
(12:05) Case Study #3: collecting evidence about AI safety parasitism
(16:46) Some final thoughts on automated alignment research
The original text contained 2 footnotes which were omitted from this narration.
---
First published:
August 13th, 2026
Source:
https://www.lesswrong.com/posts/myAhB5qyAHyXRv6KJ/automated-alignment-runs-are-hard-to-study
---
Narrated by TYPE III AUDIO.
---
Listen Now
Love PodBriefly?
If you like Podbriefly.com, please consider donating to support the ongoing development.
Support Us
