Episode Details

Back to Episodes

“An operationalization of opaque serial depth” by ryan_greenblatt, frisby, Alek Westover, Lukas Finnveden, Alexa Pan, Julian Stastny

Published 1 week, 2 days ago
Description

Currently, chain-of-thought (CoT) is a valuable tool for overseeing AI models. However, some architectural shifts could significantly reduce CoT monitorability. We have recently proposed that AI companies should transparently share ​​information about the degree to which their architectures may allow for latent reasoning and communication. To assist with this proposal, this document operationalizes a measure that serves as a proxy for the amount of unverbalized serial cognition a model can perform. Our measure is a specific instantiation of the notion of “opaque serial depth”, originally defined in a recent paper from GDM (Brown-Cohen et al, 2026).

To measure the opaque serial depth of a computation, Brown-Cohen et al. propose measuring the longest path in the computational graph which doesn’t pass through some form of “interpretable bottleneck”. Centrally, if one considers CoT tokens as “interpretable” but transformer hidden states as “non-interpretable”, then the opaque serial depth of a standard transformer is proportional to the number of layers.

Our main contribution in this document is a particular standard for what counts as an “interpretable bottleneck”. Roughly speaking, we want to consider nodes to be “interpretable bottlenecks” if they output text (as opposed to latent states), and were initialized from a pre-training [...]

---

Outline:

(05:07) Definition of natural-language-rooted nodes

(13:02) Definition of NLS depth

(19:34) NLS depth tracks concerning architecture changes

(22:43) Conclusion

(23:16) Appendix A: Applying NLS depth to plausible architectures

(23:37) Examples that don't increase NLS depth

(29:40) Examples that moderately increase NLS depth

(31:05) Examples that substantially increase NLS depth

(41:40) NLS depth for non-general-purpose models

(44:30) Appendix B: Sensible alternative operationalizations

(44:44) Other requirements for what can count as an interpretable bottleneck

(45:01) Require information bottlenecks...

(50:09) Forbid backpropagation through tokens...

(51:36) Require CoT to stay legible...

(53:35) Require paraphrase invariance...

(57:48) Other modifications

(58:02) Evaluate circuit depth at a fixed context length

(58:59) Measure opaque capabilities as opposed to circuit depth

(01:00:51) Allow opaque loops over long timescales

(01:01:34) Appendix C: Systems with high opaque serial reasoning capabilities are likely less monitorable

(01:04:29) Appendix D: Worked example of bounding depth

(01:11:36) Appendix E: NLS depth scaling is very slow for the classic transformer architecture

(01:13:38) Appendix F: Maximum FLOP of an opaque system

(01:15:39) Appendix G: Issues with low-FLOP serially intense computations

The original text contained 32 footnotes which were omitted from this narration.

---

First published:
September 10th, 2026

Source:
https://www.lesswrong.com/posts/x8BvtWxtoajBGHS3g/an-operationalization-of-opaque-serial-depth

---

Narrated by TYPE III AUDIO.

---