Episode Details

Back to Episodes

“Yet another concerning result on Astra’s no-CoT capabilities” by Christine Corry

Published 5 days, 16 hours ago
Description

This is a research update for an on-going replication of no-CoT evals done as part of the Second Look Fellowship. In following posts, we will run more comprehensive replications of previous work and release open source tooling for no-CoT eval elicitation. Code can be found here.

tl;dr

  • We replicate experiments from Greenblatt 2025 and Greenblatt 2026 on GPT-6-Astra, on the same items and protocol as our previous update on Fable 5, Opus 5, Opus 4.5, and GPT-5.6-Sol, plus Gemini 3.1 Pro, Kimi k3, and Fable 5.1.
  • We find that Astra is a qualitative jump in no-CoT capabilities over all datasets.
    • 4-hop questions: Astra achieves 31% at baseline, where every other model tested scores at 1-3%
    • 3-hop questions: 70% against previous best of 22% (Gemini 3.1 Pro)
    • Neel Nanda and Rohan Subramani report the same jump independently. Our work qualitatively replicates these results.
  • Astra sees more uplift from filler tokens and repeats than previous models
    • 4-hop performance is doubled from baseline (31%) to peak filler condition (63% at )
    • 3-hop accuracy jumps from 70% to 85%
      • Filler tokens and problem repeats raise accuracy monotonically across the full range we tested
    • Dylan Xu, SebastianP, & Alek [...]

---

Outline:

(00:29) tl;dr

(02:54) Background

(03:51) Previous work

(04:49) Datasets

(06:17) Evaluation design

(07:41) Eliciting no-CoT

(08:00) Results

(08:03) 4-Hop

(08:31) Utilization of filler tokens / problem repeats

(10:23) Per-dataset results

(10:48) Per-model profiles

(11:00) Discussion

(12:59) Related work

The original text contained 6 footnotes which were omitted from this narration.

---

First published:
September 13th, 2026

Source:
https://www.lesswrong.com/posts/tz5WvDouXKbiWJG8B/yet-another-concerning-result-on-astra-s-no-cot-capabilities

---

Narrated by TYPE III AUDIO.

---