Episode Details
Back to EpisodesCLAP: Converting Vision-Language Models into Vision-Language-Action Models via Language-Prompted Actions
Published 2 months, 3 weeks ago
Description
Converts any pretrained vision-language model into a vision-language-action model with zero architectural changes by prepending natural-language action descriptions; reaches 90.8% success on LIBERO with a 2B model after less than 6 hours of post-training.