Episode Details
Back to Episodes“Deep recurrent models are less robustly CoT-monitorable than normal CoT models in a toy setting” by Nick Kuhn, Alek Westover
Description
We use RL to teach a deep recurrent model and a normal CoT model to solve a math problem while hiding from a CoT monitor which of two possible problems it is solving. The deep recurrent model moves its reasoning into latents (thereby evading the CoT monitor) within 40 steps of RL; the CoT control is much worse at learning to confuse the LLM monitor. Note: in the above graph, “monitorability” is how much better than chance the CoT monitor is at figuring out which problem the model is solving. This was our simplest setting; similar settings show similar results with somewhat smaller differences between the CoT model and the deep recurrent model, see Results.
Some proposed architectural changes add a latent channel in addition to the normal text channel for persisting state between forward passes. These parallel latents architectures (like T2MLR and full-bandwidth transformers) could theoretically perform all their reasoning in latents, rather than text. We’ve speculated that these architectures will therefore be harder to oversee. This post empirically investigates whether these models can learn to reason without their CoT (which would likely make them harder to oversee) in a simplified setting.
We compare the performance of a [...]
---
Outline:
(03:49) A formal description of the parallel latents architecture
(05:00) Experiment 1
(05:50) Tasks
(09:09) Stage 1: SFT
(11:53) Stage 2: RL
(13:06) Results
(20:50) Discussion
(21:55) Task
(22:50) Pressure to solve the task
(23:32) Oversight method
(24:25) Model
(25:53) Obfuscation pressure
(28:04) Suggestions for future work
(29:58) Conclusion
(30:26) Acknowledgements
(30:49) Appendix
(30:52) Experiment 2: Output Spillover
(32:22) Result
(33:25) Experiment 3: Training pressure variations
(37:09) en-US-AvaMultilingualNeural__ When starting with low task performance, rejection filtering on highly monitorable trajectories can lead to increased obfuscation.
(37:20) en-US-AvaMultilingualNeural__ When initial performance on the task is high, rejection filtering does not exert significant pressure on monitorability.
The original text contained 3 footnotes which were omitted from this narration.
---
First published:
September 17th, 2026
---
Narrated by TYPE III AUDIO.
---
Images from the article:
Listen Now
Love PodBriefly?
If you like Podbriefly.com, please consider donating to support the ongoing development.
Support Us