Podcast Episodes
Back to SearchLATENT: Teaching a Humanoid to Play Tennis from Imperfect Data
Introduces a three-stage pipeline that extracts a latent action space from noisy, low-quality human motion capture, then trains a high-level RL polic…
4 months, 2 weeks ago
CollabVR: Collaborative Video Reasoning with Vision-Language and Video Generation Models
Closed-loop framework coupling Vision-Language Models with Video Generation Models at step-level granularity. Mitigates long-horizon drift and mid-cl…
4 months, 2 weeks ago
World Action Models: The Next Frontier in Embodied AI
First systematic survey defining World Action Models (WAMs) as embodied foundation models that jointly predict future states and generate actions. Co…
4 months, 2 weeks ago
Training a Whole-Body Control Foundation Model
Describes end-to-end learning of a foundation model for adaptive whole-body humanoid control via massive simulation variation. Combines proprioceptiv…
4 months, 2 weeks ago
DexJoCo: A Unified Benchmark for Task-Oriented Dexterous Manipulation
Releases an open-source MuJoCo-based benchmark with 11 dexterous tasks, low-cost teleoperation hardware, and 1.1K human demonstrations. Designed to e…
4 months, 2 weeks ago
MMSkills: Building Multimodal Skill Libraries for Visual Agents
Skill library, demonstrations, and dataset for multi-modal robotic skill learning and manipulation tasks.
4 months, 2 weeks ago
PhysBrain 1.0 VLA (TwinBrainVLA): Dual-Brain Vision-Language-Action with Physics-Grounded Learning
Introduces dual-brain fusion Vision-Language-Action model with LangForce physics-grounded training methodology.
4 months, 2 weeks ago
MolmoAct2-LIBERO: An Open Vision-Language-Action Model for Robotics
Vision-Language-Action (VLA) model fine-tuned on the merged LIBERO robotics dataset (1,693 episodes, 273k+ frames) achieving 98.25% success rate on m…
4 months, 2 weeks ago
SANA-WM: Efficient Minute-Scale World Modeling with Hybrid Diffusion Transformers
A 2.6B-parameter open-source world model that generates coherent 720p, minute-long videos with precise 6-DoF camera control on a single GPU using a H…
4 months, 2 weeks ago
WildClawBench: A Real-World, Long-Horizon Benchmark for AI Agents
New benchmark and dataset for robotic manipulation in unconstrained 'wild' environments. Includes standardized containers, leaderboards, and evaluati…
4 months, 2 weeks ago