Episode Details

Back to Episodes

“Foundation Models for Oversight” by jsteinhardt

Published 3 days, 15 hours ago
Description

Cross-posted from the Transluce blog.

To oversee an AI model, we'd ideally like to ask questions such as:

  • What are important situations where the model sandbags?
  • Does the model have an objective it wouldn't admit to if asked directly?
  • Does the model treat a user differently once it infers something about their identity, and along what axis?
  • Is the model's chain of thought load-bearing, or is it a post-hoc rationalization of an answer that was already settled on?
  • Is it reward hacking on this input, or actually trying to solve the task?

It would be great if we had an oversight assistant that could answer these questions. We'd want it to do three things: help us formalize the question as a testable empirical criterion; produce data that satisfies that criterion; and do so in a way we can justifiably trust.

To get such an assistant, we lay out a vision for building a foundation model for oversight: an AI system mid-trained (or pre-trained) on a large, diverse corpus of experiments on a given "subject model", RLVR'd on a large number of verified oversight tasks, and then fine-tuned to answer [...]

---

Outline:

(05:55) Conceptual preliminaries

(05:59) Pythonic world models

(07:21) Oversight as Inference

(10:42) Reducing oversight to autoregressive prediction

(11:45) Engineering Scale-up

(11:49) Generating supervised oversight data for mid-training

(13:35) Step 1: Sampling diverse experiments

(15:57) Step 2: Sampling diverse experiment inputs

(18:05) Step 3: Featurizing as token sequences

(20:17) Calling our shots: staged de-risking

(23:49) Stage 1: Individual Tasks

(27:14) Stage 2: Cross-Task Transfer

(28:41) Stage 3: Zero-Shot Abilities

(30:07) Finishing Touches

(30:11) RLVR

(33:55) Generalizing RLVR to Other Tasks

(35:27) Example Trajectory

(36:32) De-risking RLVR

(36:57) Stage 4: RLVR is competitive with evolution for elicitation

(37:58) Stage 5: Cross-Task Transfer for RLVR

(38:52) Post-training

(41:09) Appendix

(41:12) Further Testing the Oversight-as-Inference Hypothesis

The original text contained 8 footnotes which were omitted from this narration.

---

First published:
July 28th, 2026

Source:
https://www.lesswrong.com/posts/AqdZKyoRmN6EFCzib/foundation-models-for-oversight

---

Narrated by TYPE III AUDIO.

---