Podcast Episodes
Back to SearchHoMMI: Learning Whole-Body Mobile Manipulation from Human Demonstrations
Whole-Body Mobile Manipulation Interface (HoMMI) that learns bimanual and whole-body manipulation, long-horizon navigation, and active perception dir…
6 months, 1 week ago
VEGA-3D: Imagining 3D Worlds with Video Diffusion to Teach LLMs Spatial Reasoning
Plug-and-play framework that teaches multimodal LLMs spatial reasoning by extracting implicit 3D priors from video diffusion models, supporting geome…
6 months, 1 week ago
TurboQuant: Redefining AI Efficiency with Extreme Compression
This episode explores TurboQuant, a revolutionary set of quantization algorithms from Google Research that redefines AI efficiency through extreme co…
6 months, 1 week ago
DexWM: Learning Dexterous Object Manipulation from Human Videos
Dataset of robot trajectories designed for training world models that learn dexterous hand-object interactions from human videos, released on Hugging…
6 months, 1 week ago
FlashAttention-3: Fast & Accurate Attention with Asynchrony & Low-Precision
Major efficiency leap for Transformer attention mechanisms, enabling faster training/inference on long sequences with low-precision compute.
6 months, 1 week ago
When AI Trains on Its Own Output: The Model Collapse Problem
Warns of "model collapse" in LLMs trained on synthetic data from prior models, urging preservation of human-generated data. One of 2024's most influe…
6 months, 1 week ago
MolmoBot: A Vision-Language Model for Zero-Shot Robot Manipulation
Vision-language model (VLM) for zero-shot robot manipulation, trained entirely in simulation without real-world data; achieves 79.2% success rate on …
6 months, 1 week ago
LeWorldModel: Stable End-to-End JEPA from Pixels
A stable end-to-end Joint Embedding Predictive Architecture (JEPA) trained directly from pixels that enables robust world modeling for embodied AI sy…
6 months, 1 week ago
EgoVerse: An Egocentric Data Ecosystem for Scaling Robot Learning
Ecosystem with over 1300 hours of egocentric human video data spanning 240 scenes and 2000+ tasks, designed for scalable robot policy training via be…
6 months, 1 week ago
HSImul3R: Physics-Driven Reconstruction of Human–Scene Interactions
Physics-in-the-loop bi-directional optimization pipeline reconstructing stable, simulation-ready 3D human-scene interactions from casual videos, depl…
6 months, 1 week ago