Episode Details

Back to Episodes

mimic-video: Video-Action Models for Robot Learning

Published 2 months, 4 weeks ago
Description
Introduces Video-Action Models (VAMs) that replace standard VLMs in VLAs with a pretrained internet-scale video backbone (Cosmos-Predict2) combined with a flow-matching action decoder. Achieves approximately 10× better sample efficiency on real-world pick-and-place tasks.
Listen Now

Love PodBriefly?

If you like Podbriefly.com, please consider donating to support the ongoing development.

Support Us