Episode Details

Back to Episodes

“Quadrillion Param Costs: KV Cache, Context Length, Frontier Margins” by Vladimir_Nesov

Published 4 days, 18 hours ago
Description

The models of 2028-2031 get much bigger than the models of 2026, going from 10T total params in 2026 to maybe 240 trillion params in 2028 and then 1.4 quadrillion params in 2031, as I estimate in the previous post from HBM bandwidth/capacity, scale-up system size, pretraining compute, and scaling laws. Yet as I show in this post, if the 240T 2028 model is priced at $14/$70 per 1M input/output tokens (1.4x the API price of Mythos 5), it's going to have a 70% gross margin, and the same holds for the 1,400T 2031 model when priced at $30/$150. Cutting the price in half to $15/$75 per 1M input/output tokens lowers the gross margin to 40%, which seems painful but survivable. Going in the other direction, doubling the price to $60/$300 allows serving requests with up to 3 million tokens of context at the same 70% gross margin.

These prices rest on token costs that I calculate from first principles in this post, using estimates of future hardware specs and costs. I link the 10T 2026 model to Mythos 5 to compare with its actual API prices, also performing the calculations for my guess about Opus 4.8, and [...]

---

Outline:

(03:27) A Scaling Law for KV Cache

(07:34) Token Cost in FLOPs and Bytes

(14:31) Cost Anchors for 2025-2026

(19:35) Frontier Margins in 2026

(24:37) Chip-Time Cost of Tokens in 2028-2031

(33:48) Context Lengths and Prices in 2028-2031

The original text contained 14 footnotes which were omitted from this narration.

---

First published:
July 27th, 2026

Source:
https://www.lesswrong.com/posts/Rk6FbkDFFm8ciqefv/quadrillion-param-costs-kv-cache-context-length-frontier

---

Narrated by TYPE III AUDIO.

Listen Now

Love PodBriefly?

If you like Podbriefly.com, please consider donating to support the ongoing development.

Support Us