Episode Details

Back to Episodes

“Pretraining data, not verifiability, is why LLMs are especially good at math (and coding)” by Steven Byrnes

Published 21 hours ago
Description

Follow-up to: “LLMs are (still) mostly powered by imitative learning, not RL”

A common take I’ve been hearing is: “LLMs are especially good at math because math is easy to verify”. But that story doesn’t make much sense.

  • For one thing, “easy to verify” only matters for the RL part of LLM training pipelines, and the leading LLM companies have said that they spend very little effort on RL-for-math.
  • Worse, to the extent that the companies are doing RL-for-math, it's RLAIF, not RLVR. So really, the phrase “math is easy to verify” amounts to “LLMs are very good at judging math arguments”. But that's begging the question! Why are pretrained LLMs so much better at judging math arguments than judging, say, fiction writing? We still need an answer.

So here's a different theory, in the framework of my earlier post “LLMs are (still) mostly powered by imitative learning, not RL”:

LLMs are especially good at math because almost everything in the math literature is correct. Read a random sentence in a random math paper in the research math literature, and you can be >99% confident that the sentence is true. So if LLMs do what they do best—imitative [...]

The original text contained 3 footnotes which were omitted from this narration.

---

First published:
September 18th, 2026

Source:
https://www.lesswrong.com/posts/xvdngZAqFZfek7KGH/pretraining-data-not-verifiability-is-why-llms-are

---

Narrated by TYPE III AUDIO.

Listen Now

Love PodBriefly?

If you like Podbriefly.com, please consider donating to support the ongoing development.

Support Us