Episode Details

Back to Episodes

“Automated alignment runs are hard to study!” by Alejandro Aristizabal, draganover, Aleksandr Bowkis, Cameron Holmes

Published 1 month ago
Description

TL;DR: This post presents three case studies of automated alignment research runs at Arcadia Impact. We use these case studies to emphasise the following takeaways:

  1. It is hard to parse auto-research runs! Each run produces a couple of hundred pull requests of jargon-dense agent output. When researchers look through these logs, we find that they often come away with biased/incorrect impressions.
  2. When told to raise the score on a task, the models will sometimes brazenly cheat. It seems difficult to predict when this will happen vs. when the run will go smoothly.
  3. Hillclimbing metrics are often off-target from the spirit of an alignment task. I.e., when we use metrics as proxies for our alignment questions, we find that the models will often misunderstand the spirit of the task. This can lead to unpredictable behaviour.
  4. The runs are surprisingly reproducible. Even though a run could unfold in vastly different ways, we find that independent reruns converge on the same strategies and the same failure modes.

Models’ research capabilities are advancing quickly. If alignment is to keep pace, we may need to automate alignment research and do so responsibly. This makes it important that we have the tools to inspect [...]

---

Outline:

(04:13) Methods for analysing runs

(06:12) Case Study #1: learning synthetic concepts

(09:23) Case Study #2: training robust backdoors

(12:05) Case Study #3: collecting evidence about AI safety parasitism

(16:46) Some final thoughts on automated alignment research

The original text contained 2 footnotes which were omitted from this narration.

---

First published:
August 13th, 2026

Source:
https://www.lesswrong.com/posts/myAhB5qyAHyXRv6KJ/automated-alignment-runs-are-hard-to-study

---

Narrated by TYPE III AUDIO.

---

Images from the article:

ARCH schematic. Agents are given a task composed of a public and a held-out evaluation mechanism. A fleet of workers submits PRs that get scored against the held-out evaluation.
This plot shows the metric progression over the course of the auto-alignment run. Each dot is a single PR.
Goal alignment score on 5 workers on the
Listen Now