Podcast Episodes
Back to SearchOne Model to See, Plan, and Act: Introducing EO-1 for Embodied AI
A 3B parameter unified decoder-only transformer that interleaves vision, text, and action tokens for perception, planning, reasoning, and control in …
2 months, 1 week ago
Video-Action Models for Robot Learning
Introduces Video-Action Models (VAMs) that leverage pretrained internet-scale video models such as Cosmos-Predict2 as backbones instead of VLMs, pair…
2 months, 1 week ago
FastDSAC: Unlocking the Potential of Maximum Entropy RL in High-Dimensional Humanoid Control
Introduces FastDSAC, a method that scales maximum entropy reinforcement learning to high-dimensional humanoid control tasks, improving sample efficie…
2 months, 1 week ago
A Minimalist Retargeting-Guided Reinforcement Learning Recipe for Dexterous Manipulation
Proposes a minimalist pipeline using retargeting-guided RL for dexterous manipulation in humanoid robots. Focuses on simplifying the training recipe …
2 months, 1 week ago
DexNDM: Closing the Reality Gap for Dexterous In-Hand Rotation
DexNDM bridges the sim-to-real gap for stable in-hand rotation of complex objects, enabling learning from biased real-world data without requiring an…
2 months, 2 weeks ago
Learning Multi-Modal Trajectory Policies for Data-Efficient Robotic Manipulation
Proposes learning multi-modal trajectory policies to improve data efficiency in robotic manipulation tasks.
2 months, 2 weeks ago
ACE: Agentic Control for Embodied Manipulation via Zero-shot Workflow Reasoning
A zero-shot workflow reasoning approach for agentic control of embodied manipulation, enabling robots to perform complex tasks without task-specific …
2 months, 2 weeks ago
Let the Body Follow: Coupled Egocentric Control for Whole-Body Robot Teleoperation
A coupled egocentric control approach for whole-body robot teleoperation that coordinates body and limb movements from an egocentric perspective, sub…
2 months, 2 weeks ago
GigaWorld-Policy-0.5: An Efficient Mixture-of-Transformers Policy for Real-Time Robot Control
A Mixture-of-Transformers robot policy that achieves 85 ms inference on RTX 4090 by separating visual dynamics from action generation. Demonstrates e…
2 months, 2 weeks ago
Xiaomi-Robotics-1: A Scalable Vision-Language-Action Foundation Model for Mobile Manipulation
A scalable vision-language-action (VLA) foundation model pretrained on over 100k hours of real-world manipulation trajectories, enabling out-of-the-b…
2 months, 2 weeks ago