Episode Details

Back to Episodes
Everything Not Nailed Down Is Training Data Now -- AI Brief August 18

Everything Not Nailed Down Is Training Data Now -- AI Brief August 18

Season 2026 Episode 818 Published 1 month, 3 weeks ago
Description

Good day, humans. Most of today's brief is about what gets fed into the machine and who agreed to it. Amazon is buying rare out-of-print books and slicing the spines off to scan them. Google paid $10 million for a bankrupt airline's emails. Meanwhile Alibaba gave away a 2.4-trillion-parameter model, OpenAI switched on a million-token window it had been rationing, and a dev-tools CEO made the case that nobody is really reading the code anyway.

Amazon Is Guillotining Rare Books for Training Data

404 Media

What happened: 404 Media hid an Apple AirTag inside a 1,000-book bulk order and followed it from California to an Amazon warehouse in Las Vegas called LAS8. Workers there describe a unit named VGT3 whose job is to slice the bindings off books, scan the pages, and destroy the originals. Amazon confirmed it buys books “to help develop and improve the products and services our customers use,” and declined to say how many it has destroyed or how many such sites it runs.

Why it matters: Every model needs text it hasn't already eaten, and the open web is picked clean. Out-of-print books are the last big reservoir of writing that was never posted online and was written before 2022 — which makes it verifiably human. The problem is that for a lot of these titles, the copy going through the blade is one of the few left anywhere.

What everyone's saying: TechCrunch went straight for the irony that Amazon started life as a bookstore. Others noted Amazon isn't the first here — destructive scanning has a long industrial history — and that buying a physical book you then shred is a far cleaner legal path to training data than scraping a website and arguing about it in court for three years.

My read between the lines: This is what it looks like when a company decides the words matter and the object doesn't. Which is a defensible position right up until the scan turns out to be lossy, the model gets deprecated, and the book is landfill. We spent two decades arguing about whether AI companies should pay for the text they train on. They will. They'll buy the last copy and feed it through a paper cutter.

📖 Further reading: I Make AI Versions of Myself for a Living. This One I Didn't Agree To. — the same question one layer down: what happens when the thing being ingested never got asked.

Two of today's stories are about companies paying millions of dollars for someone else's inbox. Yours is sitting right there doing nothing. Viktor is an AI agent that lives in your Slack — connect it to any of 3,000+ tools and it builds the report, ships the dashboard, writes the code, runs the campaign. Not a chatbot you have to prompt. A coworker who files. New readers get $50 off their first month. Hire Viktor →

Google Paid $10M for a Dead Airline's Inbox

Axios

What happened: An August 14 filing in the Southern District of New York bankruptcy court shows Google won an auction for Spirit Airlines' internal business data: roughly 100 million emails, 500 million Microsoft Teams messages, 30 million lines of code, pricing data from 7.2 billion competitor flights, and payroll records goi

Listen Now

Love PodBriefly?

If you like Podbriefly.com, please consider donating to support the ongoing development.

Support Us