Episode Details

Back to Episodes

"A global workspace in language models" by wesg

Published 2 weeks, 3 days ago
Description
[This is the blog post for our new paper Verbalizable Representations Form a Global Workspace in Language Models
Readers might also be interested in: the Public commentary, Github and Neuronpedia]







As you read this sentence, circuits in your brain are adjusting your posture, controlling your breathing, and transforming lines and curves on the screen into recognizable words. Most of this processing is invisible to you. But some of what takes place in your brain you do have access to—an image that pops into your head, or a deliberate plan you make about where to go shopping. Neuroscientists and philosophers sometimes refer to the latter type of brain activity as “consciously accessible,” to distinguish it from all the other processing that goes on unconsciously. This activity has special properties: we can describe it, control it, and use it for deliberate reasoning, in contrast to all the automatic processing that goes on without our awareness.

In a new paper, we present evidence that a similar distinction has emerged in modern language models like Claude. We find that Claude has developed a small collection of internal neural patterns that, compared to all its other internal processing, play a [...]







---

Outline:

(06:09) How we found the J-space

[... 8 more sections]

---

First published:
July 6th, 2026

Source:
https://www.lesswrong.com/posts/3PaLrzxagpbnNtPLT/a-global-workspace-in-language-models

---



Narrated by TYPE III AUDIO.

---

Images from the article:

The J-space reveals internal thoughts that don’t appear in the model’s output.
Five functional properties of a global workspace, and stylized illustrations of experiments we use to test for them in language models.
J-lens readouts on six prompts, at various layers. In each case the lens surfaces an internal assessment or computation that appears nowhere in the text: the steps of a reasoning or math problem, the presence of a bug, recognition of an image, the function of a protein, and the suspicion that search results are fabricated.