Episode Details
Back to Episodes
Why Chatterbox Still Leads Open-Source TTS
Episode 4605
Published 1 week ago
Description
Ever wondered what actually powers the voices you hear on this show? We're doing a full teardown of Chatterbox, the open-source text-to-speech model from Resemble AI. We explore its non-autoregressive architecture built on a Llama backbone, the S3 tokenizer, and how it achieves parallel generation for massive speedups. We also dig into why, over a year later, it still beats newer models in production—thanks to stability, caching, and a thriving open-source ecosystem.