Episode Details
Back to EpisodesLook Where You Mean: Gaze2Act Gives VLAs a Live Intent Channel
Published 1 month, 4 weeks ago
Description
Vision-Language-Action (VLA) models have recently shown strong potential for robot learning by following language instructions. However, in practice, language alone is often insufficient to precisely convey human intent.