Episode Details
Back to Episodes
Why Speech-to-Text Still Fails at Its Own Name
Episode 4312
Published 4 weeks, 2 days ago
Description
When OpenAI's Whisper transcribed its own name as "Wispr," it exposed the fundamental flaw in speech-to-text: models that hear perfectly but understand nothing. This episode unpacks why homophone errors, dropped negations, and hallucinated punctuation survive even low word-error rates — and explores two competing architectural solutions. We compare the two-pass pipeline (Whisper + LLM cleanup) against unified multimodal models like GPT-4o that process audio and reasoning in a single pass. Which approach actually eliminates the need for human review? And what are the hidden failure modes of each? If you dictate more than a few hundred words a day, this episode will change how you think about voice input.