Episode Details
Back to Episodes“Categorical taboos are much better than threshold taboos: neuralese edition” by Linch
Description
I think what's going on in the “Does Astra use neuralese?” debate is that there's an important sense in which models *already* do self-communication in neuralese: between each layer in the forward pass the attention stream is already very hard to interpret, and clearly not in natural language.
Yet CoT monitorability is still a big deal and it'd be bad if all self-communication from models are no longer in natural language. So there have been two different proposed definitions of what is "true" neuralese:
- (My preferred) categorical definition: Since natural language currently gates recurrence in the standard transformer+CoT loop, having recurrence in neuralese is the natural category for whether something counts as "true" neuralese.
- The threshold definition. Total number X of serial steps before something appears in natural language. True neuralese counts as going above X. I think most technical experts who studied this issue, including many people at companies, prefer definition #2.
There are complicated technical arguments on both sides, but I think technical experts overall prefer #2 because they think it's more causally relevant (there's nothing inherently more difficult about monitoring a 128-layer model looped 8x than monitoring a 1024-layer model), have less weird edge cases [...]
The original text contained 3 footnotes which were omitted from this narration.
---
First published:
September 10th, 2026
---
Narrated by TYPE III AUDIO.
---
Images from the article:

Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.