Episode Details
Back to EpisodesGAM: Ground the Target Before You Learn the Action
Published 1 week ago
Description
GAM is a robot foundation model that conditions action chunk prediction on 3D grounded inputs — including language, 3D point clouds, and bounding boxes — alongside robot history, achieving strong results on LIBERO benchmarks and real-robot tasks. The 3D grounding mechanism advances how embodied agents spatially understand and act in their environment.