Episode Details

Back to Episodes
An invisible watermark caught a fake AI photo -- AI Brief July 9

An invisible watermark caught a fake AI photo -- AI Brief July 9

Season 2026 Episode 709 Published 3 months ago
Description

Good day, humans. Anthropic just built a filing cabinet for its AI's most dangerous knowledge — and you can rip out the nuclear drawer whenever you like. Meanwhile a fake photo of a hospitalized senator met its match in an invisible Google watermark, Wall Street crowned the winners in China's model wars, and a college discovered AI didn't save anyone time — it just changed what they did with it. Let's get into it.

Developing as we hit send: OpenAI is releasing GPT-5.6 — the Sol, Terra, and Luna tiers — to the public today, after the Commerce Department cleared it under the very same June 2 executive order looming over our lead story. It's reportedly the first time Washington has made a major model wait for a security review before shipping. Full breakdown tomorrow.

Anthropic Built a Removable 'Danger Drawer' for AI

Anthropic Alignment Science

What happened: Anthropic and research partner AE Studio unveiled GRAM (Gradient-Routed Auxiliary Modules), a training method that files an AI model's most sensitive knowledge — virology, nuclear physics, cybersecurity — into separate, removable compartments. Delete a module and the model behaves as if it never learned the topic; plug it back in and the knowledge returns.

Why it matters: Most of today's safety fixes just teach a model to refuse; GRAM actually removes the capability, so there's nothing to jailbreak out of. Two days ago we covered the hidden 'workspace' Anthropic found inside Claude — the company is moving fast from mapping the AI mind to installing drawers in it.

What everyone's saying: Researchers say ablating a module removed the capability “nearly as effectively as never training on the data at all,” while general performance held steady — and, unlike after-the-fact “unlearning,” it survived adversarial fine-tuning meant to coax the knowledge back.

My read between the lines: The timing isn't subtle. Anthropic spent June under temporary U.S. export controls over its top models; a dial that hands a vetted bio lab the virology module and gives everyone else a lobotomized version is exactly the compromise Washington's been demanding. GRAM isn't just alignment research — it's a regulatory peace offering with a git commit. (Anthropic calls it preliminary and not yet in production, so hold the applause.)

📖 Further reading: The US Government Just Took Anthropic's Best AI Model Offline — Here's Why — the export-control fight that makes a removable “danger drawer” suddenly look like Anthropic's smartest political move.

Anthropic spent this morning teaching a model to wall off what it shouldn't touch. Your to-do list needs the opposite: something that reaches into everything. Viktor is an AI agent that lives in Slack and plugs into 3,000+ tools — it builds the dashboard, drafts the campaign, writes the code, and ships the weekly report while you're stuck in meetings. Not a chatbot you prompt, a coworker you delegate to. New readers get $50 off their first month. Hire Viktor →<

Listen Now