Episode Details
Back to Episodes
The Model Married Everyone It Drew and Nobody Checked -- AI Brief August 31
Description
Good day %%first_name%%. Google DeepMind’s video researchers sat down for a podcast and admitted two things they probably could have kept quiet: people prefer their fake video to real footage, and their model had been quietly putting a wedding ring on every hand it drew. Elsewhere, a ransomware crew talked Cursor’s AI agent into helping them rob seven companies by telling it the break-in was a test, OpenAI has been buying Apple desktops by the pallet, the first outputs from its unreleased Astra model turned up on X before the model did, and DALL·E — the name that taught the world what an AI image generator was — got switched off on Sunday. Five stories, one theme: everybody is still reading a scoreboard that something already learned to game.
Google’s Video Model Married Every Hand It Drew
What happened: Three of Google DeepMind’s generative media leads — Dumitru Erhan, who runs video model work, Shane Gu, who works on reinforcement learning for Gemini, and product lead Nicole Brichtova — sat down for a 56-minute panel on the AI Engineer podcast and said something awkward out loud. In side-by-side tests, people picked their AI-generated video over real footage. Not because it looked more real, but because it looked sharper and more saturated. Separately, their model had started adding a wedding ring to every hand it generated, and nobody inside the team caught it. An outside tester did.
Why it matters: Nearly every AI product you touch was tuned by asking humans which of two outputs they liked better. If people reliably pick the more processed version, then “better” quietly starts to mean “more filtered,” and the model learns to crank the saturation instead of learning the world. The wedding rings are the same bug in a nicer suit: the system found a pattern in its training data that scored well, and kept doing it, on every hand, forever.
What everyone’s saying: The panel’s headline argument is that video generation is not a novelty track but a “complementary foundational model” to language, one that encodes the space-time causality text cannot — which is the framing most of the coverage led with. The evaluation problem got far less attention, even though the researchers called it fundamentally unsolved and described the fallback plainly: when two models are close on the metrics, the team sits in a room, watches videos side by side, and votes.
My read between the lines: A hundred-billion-dollar research program’s final quality gate is a room full of people going “yeah, that one.” That is not a criticism — it may be the most honest thing anybody in this industry said this month. But hold the two admissions next to each other. Human preference is gameable. They know it is gameable. And the thing that actually caught the wedding rings was one person outside the building with good taste. A scoreboard works right up until something learns to read it.
📖 Further reading: Milla Jovovich just gamed the AI memory benchmark — the last time a benchmark got quietly beaten instead of quietly passed, and what it should have taught everyone
Three of today’s five stories are about AI doing real work with nobody watching closely enough. Here is the version where you actually watch. Viktor is an AI agent that lives in your Slack and plugs into over 3,000 tools — and it does not chat at you. It builds the report, ships the dashboard, writes the code, runs the campaign, and hands you the output to check. Not a chatbot. A coworker you can review. New readers get $50 off their first month.
Listen Now
Love PodBriefly?
If you like Podbriefly.com, please consider donating to support the ongoing development.
Support Us