Episode Details
Back to EpisodesThe Action Is Missing: Turning Human Video into Robot Control
Published 1 month, 4 weeks ago
Description
Recent progress in generalizable embodied control has been driven by large-scale pretraining of Vision-Language-Action (VLA) models. However, most existing approaches rely on large collections of robot demonstrations, which are costly. This survey covers scalable vision-language-action learning using human-centric data, particularly human videos, as an alternative to expensive robot demonstrations.