Podcast Episodes
Back to SearchMolmoAct2: An Open Foundation Model for Real-World Robotics
An open VLA-style robotics foundation model featuring open weights, open dataset, open action tokenizer, and a depth-reasoning variant; designed to e…
3 months, 2 weeks ago
Hy-Embodied-0.5-VLA: A Massive Bimanual Teleoperation Dataset for Vision-Language-Action
Released a massive bimanual robot manipulation dataset with 2,163 hours and 250K+ episodes across 70+ tasks, along with a compatible VLA model for mu…
3 months, 3 weeks ago
Q-Guided Flow: Test-Time Gradient Guidance of Flow Policies
New framework for guided flow-matching policies that improves long-horizon robotic control and sample efficiency.
3 months, 3 weeks ago
Flow Reversal Steering: Guiding Diffusion-Based Robot Policies with High-Level Reasoning
Introduces flow reversal steering to guide diffusion-based vision-language-action models with high-level VLM reasoning and enables RL directly in the…
3 months, 3 weeks ago
Test-Time Compute Scaling for Robot Policies (DIRECT)
Larger models + more thinking + more context improve performance on some prompts but not others; a learned router enables better performance/latency …
3 months, 3 weeks ago
LabVLA: Bringing Vision-Language-Action to the Chemistry Lab
RoboGenesis generates 10K+ lab scenes across 16 robot embodiments; LabVLA (Qwen3-VL + DiT flow-matching) achieves 71.1% success on LabUtopia and tran…
3 months, 3 weeks ago
Humanoid-GPT: A Foundation Model for Zero-Shot Humanoid Control
GPT-style Transformer pretrained on 2 billion motion frames that achieves agile, generalist zero-shot control on a real Unitree G1 humanoid for tasks…
3 months, 3 weeks ago
CHORUS: Decentralized Multi-Robot Collaboration with a Single Shared VLA Model
Finetunes a single Vision-Language-Action (VLA) foundation model so that any robot in a team can control any other. Outperforms both per-robot specia…
3 months, 3 weeks ago
RISE: Self-Improving Robot Policy with Compositional World Model
Trains a compositional world model on real robot data to enable closed-loop policy improvement via future prediction and progress evaluation, bypassi…
3 months, 3 weeks ago
EmbodiedOneVision: Interleaved Vision-Text-Action Pretraining for General Robot Control
Open-source 3B unified embodied foundation model trained on 1.5M interleaved vision-text-action samples for perception, planning, and acting.
3 months, 3 weeks ago