Episode Details
Back to EpisodesVideo-Action Models for Robot Learning
Published 2 months, 1 week ago
Description
Introduces Video-Action Models (VAMs) that leverage pretrained internet-scale video models such as Cosmos-Predict2 as backbones instead of VLMs, paired with a flow-matching action decoder. Claims approximately 10x sample efficiency gains over standard vision-language-action models on real-world pick-and-place tasks.