Episode Details
Back to EpisodesHARP-VLA: Align the Eyes Before the Actions
Published 2 months ago
Description
Learning generalizable vision-language-action (VLA) models from large-scale human videos is promising but challenging due to cross-embodiment discrepancies in both visual observations and executable actions.