Episode Details
Back to Episodes“Almost nobody is funded to figure out what work would solve alignment” by Seth Herd
Description
Solving alignment would be easier if we worked out what problems we actually need to solve. This could be called the alignment meta-problem. Work on this problem is rarely directly funded. More focused work on it should let us use our limited time and funding more efficiently.
The diagram implies narrowing alignment work, but I expect meta-problem work to also identify high-payoff "fringe" approaches.
If we're driving toward a cliff, maybe we should buy better headlights.
All too often we're doing work that merely sounds or feels good, and optimizing less than we could for work that drives most efficiently toward success. Some of this is inevitable and some of it is useful, but we could do more to light the path ahead.
Most researchers agree that mech interp, refining and improving alignment training, control, theory, and miscellaneous techniques like confession are useful for solving alignment. Working toward regulation and slowdown/pause is also commonly considered useful in the governance space, and spreading awareness of alignment risks is pretty obviously useful for accelerating and enabling all of this work. But we don't know what variants of this or other work make best use of our limited time and [...]
---
Outline:
(03:03) Why not to fund more work on the meta-problem
(03:50) Arguments in favor, compressed
---
First published:
September 4th, 2026
---
Narrated by TYPE III AUDIO.
---
Images from the article:

Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.