Episode Details

Back to Episodes

“Toy Models of Initialisation Effects on RL Dynamics” by Edward James Young, lennie

Published 2 weeks, 4 days ago
Description

This is a follow-up to two posts Geodesic released last week on our current research direction. The code for generating the figures can be found at this GitHub repository.

In our previous post, we outlined Geodesic's focus on what we term the pre-RL alignment checkpoint of models -- the alignment-relevant properties of a model conveyed by pretraining, midtraining, and warm-start SFT, going into heavy RL post-training. In this post, we analyse a toy model of RL learning dynamics, with a particular focus on the effect of initialisations, to illustrate some of the ideas that we introduced.

There are three main ideas we'll use our toy model to illustrate. For a more detailed discussion of these ideas in the context of frontier post-training runs, see the previous post.

  • Rich-get-richer dynamics. The solution that the model learns can depend importantly on the initial strategies into which it explores.
  • Underspecified behaviours. When the reward function doesn't depend on an aspect of a model's behaviour -- such as its emotional state while performing a task, or a belief that its reality is simulated -- those behaviours might be primarily determined by the pre-RL checkpoint.
  • Underspecification of model cognition. As a special case of the [...]

---

Outline:

(01:51) Mathematical preliminaries

(02:41) RL displays rich-get-richer dynamics

(05:31) Underspecified behaviours

(06:07) Where do these dynamics come from?

(07:32) Unequal rewards

(11:27) Learning when the reward underspecifies cognition

(16:07) Discussion

(17:03) Extensions

(20:32) Author contributions

(20:59) Appendices

The original text contained 6 footnotes which were omitted from this narration.

---

First published:
July 14th, 2026

Source:
https://www.lesswrong.com/posts/72AAjXAxS7Pow9Fie/toy-models-of-initialisation-effects-on-rl-dynamics

---

Narrated by TYPE III AUDIO.

---

Images from the article:

Figure 1. Deterministic full-gradient dynamics for rewards. Left: the full-gradient update at each policy on the simplex, arrows scaled by magnitude. Right: five full-gradient trajectories starting at (dashed line) with run at for steps and coloured light-to-dark by step.
Figure 2. Twenty finite-sample RLOO rollouts for, all starting at, and coloured light-to-dark by step. Below each panel, the density of the final policy's horizontal position -- from vertex (left) to vertex (right) -- estimated over rollouts on a shared vertical scale. Left:,, steps. Right:,, steps.
Listen Now

Love PodBriefly?

If you like Podbriefly.com, please consider donating to support the ongoing development.

Support Us