Episode Details

Back to Episodes

“Measuring Activation Control in LLMs” by Marek Kowalski, Joshua Fonseca Rivera, Uzay Macar, David Africa

Published 1 month ago
Description

TL;DR

  • Inspired by the introspective awareness and CoT controllability papers, we made a benchmark to measure how well models can control their activations while completing a simple task.
  • We are motivated by the concern that highly introspective models could control their activations, confounding probes and other monitors, and potentially even influencing their own training.
  • We ran this on 25 open weight models ranging from 4B to 744B. We find that most language models are able to not only increase the salience of a concept in their residual stream on command, but also dial its strength up and down, including during specific intervals relative to the duration of the task. We also find that models are unable to control at which specific layer this is done.
  • Counterintuitively, we find that within five of the seven model families we tested, the newest model scores lowest. For some reason, one of the oldest and smallest models of the panel, Llama 3.1 8B, performs best.
  • It's not clear to us that newer models should have poorer control over their internal representations. More likely, where they “think” stops being the activation space, and becomes something else. We are looking for feedback (and other possible [...]

---

Outline:

(00:13) TL;DR

(02:01) Methods

(08:58) Results

(15:41) Discussion

(16:29) Acknowledgements

---

First published:
August 12th, 2026

Source:
https://www.lesswrong.com/posts/HgvwxjzgwvsEvAiBH/measuring-activation-control-in-llms

---

Narrated by TYPE III AUDIO.

---

Images from the article:

Line graphs showing
Two line graphs showing d' and dial rank ρ versus depth percentage.
Bar graph comparing
Listen Now

Love PodBriefly?

If you like Podbriefly.com, please consider donating to support the ongoing development.

Support Us