Podcast Episodes
Back to SearchEP106: Prompt Caching in Production — Lower Cost Without Stale Behavior
A practical guide to prompt caching for AI APIs: cacheable prefixes, routing consistency, measurement, invalidation, privacy, and the production chec…
2 months ago
EP105: Claude Opus 5 Is Live — A Practical Evaluation Plan for Developers
Claude Opus 5 is now available on Crazyrouter. Here is a disciplined plan for evaluating a new flagship model on quality, reliability, latency, and c…
2 months, 1 week ago
EP104: AI API Cost Guardrails — Keep Production Spend Predictable
A practical framework for AI API budgets, token limits, routing rules, alerts, quotas, and circuit breakers that prevent surprise production costs.
2 months, 1 week ago
EP103: Model Capability Contracts — Stop Silent Failures Across Providers
How to define and enforce model capability contracts for tools, JSON output, context limits, vision, streaming, retries, and safe multi-provider rout…
2 months, 1 week ago
EP102: AI API Failover Drills — Test Recovery Before Production Breaks
How to run practical AI API failover drills covering timeouts, rate limits, invalid output, provider outages, routing, idempotency, and recovery metr…
2 months, 1 week ago
EP099: Idempotency for AI APIs — Preventing Duplicate Work and Duplicate Charges
How to make AI API workflows safe under retries: idempotency keys, request state, streaming recovery, tool side effects, and billing reconciliation.
2 months, 1 week ago
EP101: The Metrics That Matter Before You Switch AI Models
A practical guide to evaluating AI models in production: accepted-result rate, latency, retries, fallback share, and cost per accepted result.
2 months, 1 week ago
EP100: AI API Gateways at 100 Episodes — What We Learned About Reliable Model Access
A 100-episode retrospective on AI API gateways: model access, pricing, routing, validation, observability, content workflows, and the reliability les…
2 months, 1 week ago
EP098: Cost per Accepted Result — The Metric AI Teams Actually Need
Why AI API pricing should be measured by cost per accepted result, including retries, validation, fallback routing, latency, and human repair.
2 months, 1 week ago
EP097: Building a Reliable AI API Gateway — Timeouts, Budgets, and Observability
A practical architecture guide for reliable AI API gateways: deadlines, retries, token budgets, streaming, validation, fallback routing, and observab…
2 months, 1 week ago