Podcast Episodes
Back to SearchHANDOFF: Humanoid Agentic Task-Space Whole-Body Control via Distilled Complementary Teachers
A new humanoid whole-body control method that distills complementary teachers for agentic task-space control on humanoids. The approach enables robus…
2 months, 3 weeks ago
EVA-Client: A Unified Framework for Real-Robot Policy Iteration
A unified open framework for real-robot policy iteration that integrates teleoperation data collection, model training, deployment, and evaluation in…
2 months, 3 weeks ago
UniVR-34B: A Vision-Only Foundation Model for Physical Tasks
First large-scale model to learn complex physical dynamics, visual reasoning, and long-horizon planning directly from visual demonstrations without t…
2 months, 3 weeks ago
CLAP: Converting Vision-Language Models into Vision-Language-Action Models via Language-Prompted Actions
Converts any pretrained vision-language model into a vision-language-action model with zero architectural changes by prepending natural-language acti…
2 months, 3 weeks ago
InternVLA-A1.5: Unifying Understanding, Latent Foresight, and Action
InternVLA-A15 is a new Vision-Language-Action (VLA) model released by InternRobotics, targeting robotic manipulation and control tasks. The model is …
2 months, 3 weeks ago
MotionWAM: Towards Foundation World Action Models for Real-Time Humanoid Loco-Manipulation
Proposes foundation world-action models targeting real-time humanoid loco-manipulation, integrating world modeling with action generation for unified…
2 months, 3 weeks ago
On-Device Diffusion Transformer Policy for Efficient Robot Manipulation
LightDP introduces a framework to accelerate Diffusion Transformer-based policies for real-time, on-device deployment in robot manipulation tasks, ac…
2 months, 3 weeks ago
Robotic World Model: Neural Dynamics for Locomotion
A neural dynamics world model paired with model-free RL for quadruped and humanoid locomotion in IsaacLab, enabling long-horizon autoregressive predi…
2 months, 3 weeks ago
LingBot-Vision: Self-Supervised ViT with Masked Boundary Modeling
A self-supervised ViT backbone pretrained with masked boundary modeling for dense spatial perception, achieving strong zero-shot results on depth est…
2 months, 3 weeks ago
LaMem-VLA: Latent-Memory-Native VLA Framework
A VLA framework that curates experience into short- and long-term memory vaults, condenses them to latent tokens, and integrates directly into VLA re…
2 months, 3 weeks ago