Episode Details
Back to Episodes
How AI Training Data Gets Filtered (and Exploited)
Episode 5117
Published 1 month ago
Description
When a model says something toxic or wrong, what actually kept it out of the training data? This episode maps the full content filtering pipeline — from Common Crawl's 250 billion raw pages to the final curated dataset. We break down the six stages: language ID, deduplication, quality scoring, toxicity filters, and curation, plus the tools like FineWeb, Dolma, and Datatrove that do the work. Then we explore why persistent pre-training poisoning attacks succeed even against aggressive filtering — and why the human annotators meant to catch bad content might be the weakest link.
Episode #920371 — open it directly at myweirdprompts.com/920371