Episode Details
Back to Episodes
Testing the Unpredictable: QA for Agentic AI
Episode 4444
Published 2 weeks, 6 days ago
Description
Traditional software testing assumes determinism — run the same test, get the same result. Agentic AI shatters that assumption. This episode maps the emerging QA landscape for probabilistic systems: from golden datasets and LLM-as-judge to trajectory evaluation and adversarial prompting suites like Garak. We explore what carries over from traditional testing, what requires entirely new methodologies, and what a sane minimum testing stack looks like for teams shipping agentic systems.