Episode Details

Back to Episodes
Everything Got Cheaper Except Your Attention -- AI Brief August 12

Everything Got Cheaper Except Your Attention -- AI Brief August 12

Season 2026 Episode 812 Published 1Β month, 4Β weeks ago
Description

Good day, humans. The floor fell out of AI pricing today, and not because anyone in San Francisco decided it should. Nvidia gave away a model that runs on the graphics card already sitting in your gaming PC. Gemini crossed a billion people. Spotify started putting badges on bands that don't exist. And one builder came back from a river vacation to find eleven terminal tabs of work nobody had asked for. Cheap and everywhere both arrived this morning. The labels are running late.

China Set the Price. America Paid It.

Source: South China Morning Post

What happened: A wave of cheap, capable Chinese models β€” DeepSeek's V4-Flash, Alibaba's Qwen line, Moonshot's Kimi K3 β€” has pushed American labs into cutting API prices to defend share. DeepSeek charges roughly $0.14 per million input tokens. OpenAI cut its lightweight GPT-5.6 Luna tier by 80%, and Anthropic shipped a near-flagship model at half the old price.

Why it matters: A million words of machine thinking now costs closer to a coffee than to an employee. If you were waiting for AI to get cheap enough to put inside your own product, that wait ended somewhere around last week.

What everyone's saying: AFP's roundup, carried by Hong Kong Free Press, got the careful analyst line: this is β€œa price competition,” not yet a price war. Forbes was blunter and called it a race to the bottom.

My read between the lines: Look at where nobody cut. Luna dropped 80%; the top-end Sol tier didn't move. The cheap tier is where the volume lives, so that's where the knife fight is. The frontier tier is where the story about needing hundreds of billions in compute lives, and that story still has to hold. These cuts aren't generosity and they aren't surrender. They're a company choosing which half of its business it's willing to lose money on.

πŸ“– Further reading: Fable 5 Costs 2x Opus β€” and Using It Wrong Costs You More Than That β€” when the cheap tier gets this cheap, the money you waste is the money you spend sending work to the expensive one.

Story one is about the price of tokens falling. Nobody is cutting the price of the two hours a week you spend pulling numbers into a deck nobody reads. Viktor is an AI agent that lives in your Slack and connects to over 3,000 tools, and it does the work: the report, the dashboard, the code change, the campaign. Not a chatbot you ask questions. A coworker you hand things to. New readers get $50 off their first month. Hire Viktor β†’

Nvidia Put a Frontier-Class Model in Your Gaming PC

Source: CNBC

What happened: Nvidia released Nemotron 3.5 Lightning, its first open-source model since Jensen Huang started publicly defending open weights. It is a 30-billion-parameter mixture-of-experts model that wakes only 3 billion parameters per token, so it runs on one consumer graphics card. Companies can use, adapt and redistribute it without asking Nvidia. Alongside it came NeMo Switchyard, an open library that routes agent requests to the cheapest model that can handle them.

Why it matters: Yesterday we ran this under

Listen Now