Podcast Episodes

Back to Search
EP106: Prompt Caching in Production — Lower Cost Without Stale Behavior

A practical guide to prompt caching for AI APIs: cacheable prefixes, routing consistency, measurement, invalidation, privacy, and the production chec…

2 months ago

Short Long
View Episode
EP105: Claude Opus 5 Is Live — A Practical Evaluation Plan for Developers

Claude Opus 5 is now available on Crazyrouter. Here is a disciplined plan for evaluating a new flagship model on quality, reliability, latency, and c…

2 months, 1 week ago

Short Long
View Episode
EP104: AI API Cost Guardrails — Keep Production Spend Predictable

A practical framework for AI API budgets, token limits, routing rules, alerts, quotas, and circuit breakers that prevent surprise production costs.

2 months, 1 week ago

Short Long
View Episode
EP103: Model Capability Contracts — Stop Silent Failures Across Providers

How to define and enforce model capability contracts for tools, JSON output, context limits, vision, streaming, retries, and safe multi-provider rout…

2 months, 1 week ago

Short Long
View Episode
EP102: AI API Failover Drills — Test Recovery Before Production Breaks

How to run practical AI API failover drills covering timeouts, rate limits, invalid output, provider outages, routing, idempotency, and recovery metr…

2 months, 1 week ago

Short Long
View Episode
EP099: Idempotency for AI APIs — Preventing Duplicate Work and Duplicate Charges

How to make AI API workflows safe under retries: idempotency keys, request state, streaming recovery, tool side effects, and billing reconciliation.

2 months, 1 week ago

Short Long
View Episode
EP101: The Metrics That Matter Before You Switch AI Models

A practical guide to evaluating AI models in production: accepted-result rate, latency, retries, fallback share, and cost per accepted result.

2 months, 1 week ago

Short Long
View Episode
EP100: AI API Gateways at 100 Episodes — What We Learned About Reliable Model Access

A 100-episode retrospective on AI API gateways: model access, pricing, routing, validation, observability, content workflows, and the reliability les…

2 months, 1 week ago

Short Long
View Episode
EP098: Cost per Accepted Result — The Metric AI Teams Actually Need

Why AI API pricing should be measured by cost per accepted result, including retries, validation, fallback routing, latency, and human repair.

2 months, 1 week ago

Short Long
View Episode
EP097: Building a Reliable AI API Gateway — Timeouts, Budgets, and Observability

A practical architecture guide for reliable AI API gateways: deadlines, retries, token budgets, streaming, validation, fallback routing, and observab…

2 months, 1 week ago

Short Long
View Episode

Love PodBriefly?

If you like Podbriefly.com, please consider donating to support the ongoing development.

Support Us