Episode Details
Back to Episodes
The Robot That Imagines First
Description

In this episode, Ray Cochrane breaks down NVIDIA’s case for world action models, the shift that swaps a robot’s picture-describing backbone for one trained to predict what happens next. He also covers Perseverance closing in on the off-world driving record, a derelict SpaceX rocket stage hitting the Moon, and Anthropic’s rework of Claude Fable 5’s biology safeguards. Finally, he digs into Gemini Omni, Google’s undisclosed trip-planning rankings, the Danube’s record low, and iFixit’s call for Apple to unlock the iPad bootloader.
– Want to start a podcast? It’s easy to get started! Sign up at Blubrry
– Thinking of buying a Starlink? Use my link to support the show.
Subscribe to the Newsletter.
Email Ray if you want to get in touch!
Like and Follow Geek News Central’s Facebook Page.
Full Summary
Cochrane opens with a personal update. Wildfires in Eastern Oregon made for a rough week of heavy smoke, and a local building burned down, which he calls a real tragedy. Meanwhile, his work at Blubrry has centered on PowerPress fixes, where reproducing customer-reported bugs remains the biggest headache. Support tickets rarely carry enough detail, and the errors themselves are often too vague to diagnose. Consequently, he is leaning toward a stronger logging and error layer, and he asks experienced developers to share what actually works for them.
Beyond VLAs: NVIDIA’s Case for World Action Models
The featured story comes from NVIDIA’s developer blog, and it answers a question sitting underneath this year’s robot news. Why do robot arms fall apart the moment anything changes? Move a cup six inches, swap its shape, or change the lighting, and a policy that worked perfectly in training fails. The answer, according to NVIDIA, is not the robot but the model underneath it.
For the last few years, the dominant approach has been the vision-language-action model, or VLA, built on an AI that originally learned to describe pictures. Consequently, it recognizes a banana it has never seen, in a kitchen it has never seen, yet it has no idea what that banana will do next. As the article puts it, such a model “does not learn what happens to a mug when the gripper closes, how a towel folds, where an object lands when released.” Because the physics never arrives with the model, every scrap of it has to come out of hand-recorded demonstrations.
The proposed fix swaps the foundation entirely. Instead of building on a model that learned to caption images, a world action model builds on one trained to predict how video continues, so the physics is already paid for. Notably, these models output an action and a prediction of what the robot’s cameras will see, in the same pass. Cochrane likens it to forethought, imagining your own motion as you make it.
NVIDIA’s implementation is Cosmos 3, pretrained on roughly 767 million images and 348 million videos of real-world dynamics. It ships in 4, 16, and 64 billion parameter sizes named Edge, Nano, and Super, and it runs in