Episode Details
Back to EpisodesLingBot-Vision: Self-Supervised ViT with Masked Boundary Modeling
Published 2 months, 3 weeks ago
Description
A self-supervised ViT backbone pretrained with masked boundary modeling for dense spatial perception, achieving strong zero-shot results on depth estimation, segmentation, and embodied tasks.