Episode Details

Back to Episodes

“Internal State Control is a General Property of LLMs” by Finn Cairns

Published 1 day, 5 hours ago
Description

tl;dr:

  • Lindsey 2025 found models can modulate their internal states: when instructed to “think about” a concept while writing an unrelated sentence, the representation of the concept is more present than when instructed to not think about it.
  • Internal state controllability appears to be a general property of LLMs: the effect replicates in 14 open-weight models from 0.3B to 235 billion parameters (Qwen3, Gemma 3, Tulu 3) with no clear trend in the think vs. don't-think gap across scale.
  • Since controllability is present even at ≤1 billion parameters with no size trend, we suspect there is a simpler attention-tagging mechanism at play, rather than metacognition.
  • Current open weight LLMs cannot weaponize this controllability: in a sandbagging setup, the model cannot evade a deception probe when instructed to suppress its signal.

There's an interactive widget here in the post.

Figure 1: Cosine similarity between the concept vector and residual stream at each layer averaged over tokens of the prefilled assistant response, under the think and don’t think prompts, for the Qwen3 model family. The gray region is a baseline of 95% CI of the cosine similarity of unrelated concept vectors, and the shaded bands are ±1 SEM.

This replication was [...]

---

Outline:

(00:11) tl;dr:

(01:45) Background

(03:25) Extending to additional model families

(04:18) Internal State Control Is a General Property

(05:22) Model Size Does Not Influence Controllability

(06:37) Silent Representations

(09:10) Prompted Model Organisms Cannot Evade Probes

(10:07) Why This Is General

The original text contained 3 footnotes which were omitted from this narration.

---

First published:
July 30th, 2026

Source:
https://www.lesswrong.com/posts/Dvqmgfeu2KDF7uMkx/internal-state-control-is-a-general-property-of-llms

---

Narrated by TYPE III AUDIO.

---