Podcast Episodes
Back to SearchEP098: Cost per Accepted Result — The Metric AI Teams Actually Need
Why AI API pricing should be measured by cost per accepted result, including retries, validation, fallback routing, latency, and human repair.
2 months, 1 week ago
EP097: Building a Reliable AI API Gateway — Timeouts, Budgets, and Observability
A practical architecture guide for reliable AI API gateways: deadlines, retries, token budgets, streaming, validation, fallback routing, and observab…
2 months, 1 week ago
EP096: The Real Cost of AI API Failures — Retries, Truncation, and Routing
A practical guide to measuring AI API reliability beyond HTTP 200: incomplete outputs, retries, validation, fallback routing, and cost per accepted r…
2 months, 1 week ago
EP095: Kimi K3 vs GLM-5.2 — Why Output Completeness Beats Leaderboards
A practical benchmark-driven comparison of Kimi K3 and GLM-5.2 across university-level math, physics, and production Python, with a focus on correctn…
2 months, 2 weeks ago
EP094: Kimi K3 vs Claude Fable 5 — Choosing Models by Workflow, Not Hype
A practical comparison of Kimi K3 and Claude Fable 5 for coding, long-context work, structured output, latency, and cost-aware AI application routing.
2 months, 2 weeks ago
EP093: Fallbacks Are Product Features, Not Afterthoughts
A practical episode on designing AI API fallbacks as a first-class product feature: routing rules, compatibility checks, retry budgets, observability…
2 months, 2 weeks ago
EP092: Model Catalogs Are Production Interfaces
A practical episode on why AI model catalogs need operational discipline: exact aliases, access rules, pricing visibility, endpoint examples, depreca…
2 months, 3 weeks ago
EP091: Shipping New Claude Models Without Breaking Developer Workflows
A practical episode on launching new Claude-family models in an API gateway: model aliases, unified endpoints, token allowlists, docs language, test …
3 months ago
EP090: Video APIs Need Real Task Evidence, Not Just Model Lists
A practical episode about production video generation APIs: async task submission, polling, model availability, endpoint drift, archived video URLs, …
3 months ago
EP089: GLM-5.2 and the New Shape of Long-Horizon AI APIs
A practical episode about GLM-5.2, long-horizon coding models, 1M-token context, reasoning-token budgets, unlimited RPM promises, and what engineerin…
3 months, 1 week ago