Episode Details
Back to Episodes
Voice AI Agents: Testing, Trust, and Scale
Description
Getting a voice agent to talk is one challenge. Knowing whether it can handle real conversations is another.
Brooke Hopkins, founder of Coval, joins The Tech Trek to discuss how teams test, evaluate, and scale voice AI agents. Drawing on her experience leading evaluation infrastructure at Waymo, Brooke explains why testing voice agents shares surprising similarities with testing autonomous vehicles.
The conversation explores why older voice assistants struggled, how better reasoning models changed what's possible, and what it takes to earn users' trust.
Brooke also shares how her engineering team works with multiple AI agents, why technical interviews need to assess AI skills, and how she balances automation with human judgment as a founder.
Key Takeaways
- Voice AI evaluations must account for interruptions, background noise, and complex conversations.
- Simulations identify failures before deployment, while production QA reveals what testing missed.
- Engineers need new ways to divide work when running several AI agents in parallel.
- Hiring should assess problem solving, system design, and how candidates use AI.
Episode Highlights
00:32 How Coval tests voice agents before and after deployment
01:26 What autonomous vehicle simulation teaches us about voice AI
03:24 Why better speech recognition wasn't enough
08:00 Could voice AI change how we use screens?
14:43 How AI is changing engineering interviews
16:43 Where a founder uses AI and where human judgment matters
A Moment Worth Pulling Out
"We had the eyes and the mouth, but we didn't have the brain."
Follow The Tech Trek for more conversations about engineering, AI, and building technology companies.