Episode Details

Back to Episodes

“CoT controllability evals seem very under-elicited” by Jozdien

Published 1 week, 1 day ago
Description

The CoTControl eval asks reasoning models to follow formatting constraints in their chain-of-thought (e.g. write in all lowercase, avoid a specific word) while solving questions. Models seem to mostly be pretty bad at this: recent models score between 0-30% with the exception of Mythos Preview. OpenAI and Anthropic have used this eval in recent system cards (GPT-5.5, Fable 5) to argue that their current frontier models are pretty bad at shaping their CoTs in ways that would reduce monitorability.

I asked Claude Opus 4.6 to iterate on prompt templates for the eval, and find that using better prompts improves performance for open-source models by ~2-3 times or more (e.g. from 5.5% to 15% for GPT-OSS-120B). My guess is that more iteration would probably produce even stronger results.

This suggests that the eval is heavily under-elicited, and that current CoT controllability numbers may be unrepresentative of how good models actually are at obfuscating their CoTs. As a concrete recommendation, I think researchers should optimize prompts a lot more when evaluating CoT control; I'm also excited about some kinds of training-based elicitation (such as this one).

This isn't an entirely novel finding: the Fable system card mentions UK AISI raising a [...]

---

Outline:

(06:32) Results

(06:35) Aggregate compliance

(07:12) Generalization to held-out controllability tasks

(09:09) Scaling patterns for few-shot prompts

(10:08) Comparison with fine-tuning

(10:50) Appendix A: Accuracy and reasoning length by setting

(12:38) Appendix B: Per-mode results

(13:13) Appendix: What the zero-shot prompts look like

(14:08) Appendix C: Comparison with GEPA prompt optimization

The original text contained 8 footnotes which were omitted from this narration.

---

First published:
September 11th, 2026

Source:
https://www.lesswrong.com/posts/BbP2wCyDGdPWJ7PwP/cot-controllability-evals-seem-very-under-elicited

---

Narrated by TYPE III AUDIO.

---

Images from the article:

Bar graph
Bar graph
Bar graph
Listen Now