Episode Details
Back to EpisodesTurboVLA and the Case for Taking the LLM Out of the Servo Loop
Published 2 months ago
Description
Achieves high-speed inference for vision-language-action policies at 32 Hz on an RTX 4090 with less than 1 GB VRAM using a 0.2B parameter model, enabling real-time VLA deployment on consumer hardware.