Podcast Episodes
Back to SearchStop Making the Video Model Remember the World
A new world model architecture that decouples world-state evolution from rendering: an agent writes executable rules, a persistent engine tracks enti…
3 weeks, 2 days ago
Atlas Turns Phone Scans Into Robot Worlds—But Not Yet Into Physics
Atlas is a unified spatial intelligence model that reconstructs real scenes from 2–3 images, generates 1440p video along camera paths, and supports r…
3 weeks, 2 days ago
W²G-Net: Carrying Instance Identity from Shiny Scrap to the Gripper
Purpose This study aims to develop a robust vision-guided framework for instance-level recognition and grasp-based classification of metal waste in d…
3 weeks, 3 days ago
When Avoidance Is Not Enough: T-CARE Gives UAV Swarms a Right-of-Way
Multi-agent UAV navigation in cluttered environments is challenged by deadlock, starvation-like persistent yielding, oscillation, and unsafe congesti…
3 weeks, 3 days ago
Teaching Robots What "Better" Means
A method for preference-based learning applied to dexterous robotic manipulation, enabling robots to learn from human preferences in freeform setting…
3 weeks, 4 days ago
TANGO: Stop Treating a Humanoid Like a Disk on the Floor
TANGO directly predicts 29-DoF joint actions from RGB observations for zero-shot real-world humanoid navigation, enabling coordinated arms, torso, an…
3 weeks, 4 days ago
Stop Teaching the VLA Your Camera: Robot-Centric Pointmaps as an Action-Aligned Visual Interface
Vision-language-action (VLA) models predict robot actions from visual observations and language instructions. These actions are defined in the robot'…
3 weeks, 5 days ago
One Policy, Many Bodies: Qwen-VLA's Bid to Unify Embodied AI
Embodied intelligence is often studied through specialized models for individual tasks such as manipulation or navigation, resulting in fragmented ca…
3 weeks, 5 days ago
When the Simulator Lies About Friction: Physics as a Transfer Prior for Touch
Physics-Regularized Learning for Robust Sim-to-Real Tactile Friction Estimation In this episode of Embodied AI 101, we explore "When the Simulator Li…
3 weeks, 6 days ago
Seeing Around Corners, Acting at the Edge: REACT as a VLA Safety Co-Pilot
Edge-based multimodal sensor data fusion with Vision-Language-Action (VLA) model for real-time autonomous vehicle accident avoidance In this episode …
3 weeks, 6 days ago