Episode Details
Back to Episodesmimic-video: Video-Action Models for Robot Learning
Published 2 months, 4 weeks ago
Description
Introduces Video-Action Models (VAMs) that replace standard VLMs in VLAs with a pretrained internet-scale video backbone (Cosmos-Predict2) combined with a flow-matching action decoder. Achieves approximately 10× better sample efficiency on real-world pick-and-place tasks.