Podcast Episodes
Back to SearchLingBot-VA 2.0: Native Video-Action Foundation Model for Robot Control
A native video-action foundation model pretrained from scratch with a semantic visual-action tokenizer and foresight reasoning, enabling real-time ro…
2 months, 3 weeks ago
GigaWorld-1: Large-Scale World Model for Robot Policy Evaluation
A large-scale world model trained on 12,980 hours of data and benchmarked across 324,000+ simulated rollouts, designed for comprehensive robot policy…
2 months, 3 weeks ago
Learning Unified Force and Position Control for Legged Loco-Manipulation
Proposes a unified RL policy trained in simulation that jointly handles force and position control without force sensors, enabling compliant behavior…
2 months, 3 weeks ago
Robix: Unified Vision-Language Model for Robotic Reasoning and Planning
A single vision-language model that unifies high-level reasoning, long-horizon planning, and human-robot interaction, serving as a cognitive layer ab…
2 months, 3 weeks ago
EmbodiedOneVision (EO-1): A Unified Decoder-Only Transformer for General Robot Control
A 3B-parameter open-source unified decoder-only transformer for general robot control that jointly handles perception, planning, reasoning, and actio…
2 months, 4 weeks ago
mimic-video: Video-Action Models for Robot Learning
Introduces Video-Action Models (VAMs) that replace standard VLMs in VLAs with a pretrained internet-scale video backbone (Cosmos-Predict2) combined w…
2 months, 4 weeks ago
Lowering the Barrier: The GEM Open-Source Arm
An open-source, sub-$500 7-DOF 3D-printed robotic arm with 1.2 kg payload, wrist and head cameras, and full LeRobot integration, designed to lower th…
2 months, 4 weeks ago
RynnWorld-4D: A 4D Embodied World Model for Bimanual Robot Control
A 4D embodied world model that predicts RGB, depth, and optical flow from RGB-D input and instructions using a tri-branch diffusion architecture. Ena…
2 months, 4 weeks ago
RynnWorld-Teleop: Digital Teleoperation via World Model Rendering
A digital teleoperation system that uses hand-pose streams to drive a world model for real-time (40+ FPS) high-fidelity robot video rendering from a …
2 months, 4 weeks ago
VLA-Corrector: Adaptive Action Horizons through Latent Visual Monitoring
A lightweight plug-in that monitors latent visual dynamics in action-chunked policies, drops stale actions on drift, and replans on-the-fly achieving…
2 months, 4 weeks ago