Podcast Episodes

Back to Search
EP098: Cost per Accepted Result — The Metric AI Teams Actually Need

Why AI API pricing should be measured by cost per accepted result, including retries, validation, fallback routing, latency, and human repair.

2 months, 1 week ago

Short Long
View Episode
EP097: Building a Reliable AI API Gateway — Timeouts, Budgets, and Observability

A practical architecture guide for reliable AI API gateways: deadlines, retries, token budgets, streaming, validation, fallback routing, and observab…

2 months, 1 week ago

Short Long
View Episode
EP096: The Real Cost of AI API Failures — Retries, Truncation, and Routing

A practical guide to measuring AI API reliability beyond HTTP 200: incomplete outputs, retries, validation, fallback routing, and cost per accepted r…

2 months, 1 week ago

Short Long
View Episode
EP095: Kimi K3 vs GLM-5.2 — Why Output Completeness Beats Leaderboards

A practical benchmark-driven comparison of Kimi K3 and GLM-5.2 across university-level math, physics, and production Python, with a focus on correctn…

2 months, 2 weeks ago

Short Long
View Episode
EP094: Kimi K3 vs Claude Fable 5 — Choosing Models by Workflow, Not Hype

A practical comparison of Kimi K3 and Claude Fable 5 for coding, long-context work, structured output, latency, and cost-aware AI application routing.

2 months, 2 weeks ago

Short Long
View Episode
EP093: Fallbacks Are Product Features, Not Afterthoughts

A practical episode on designing AI API fallbacks as a first-class product feature: routing rules, compatibility checks, retry budgets, observability…

2 months, 2 weeks ago

Short Long
View Episode
EP092: Model Catalogs Are Production Interfaces

A practical episode on why AI model catalogs need operational discipline: exact aliases, access rules, pricing visibility, endpoint examples, depreca…

2 months, 3 weeks ago

Short Long
View Episode
EP091: Shipping New Claude Models Without Breaking Developer Workflows

A practical episode on launching new Claude-family models in an API gateway: model aliases, unified endpoints, token allowlists, docs language, test …

3 months ago

Short Long
View Episode
EP090: Video APIs Need Real Task Evidence, Not Just Model Lists

A practical episode about production video generation APIs: async task submission, polling, model availability, endpoint drift, archived video URLs, …

3 months ago

Short Long
View Episode
EP089: GLM-5.2 and the New Shape of Long-Horizon AI APIs

A practical episode about GLM-5.2, long-horizon coding models, 1M-token context, reasoning-token budgets, unlimited RPM promises, and what engineerin…

3 months, 1 week ago

Short Long
View Episode

Love PodBriefly?

If you like Podbriefly.com, please consider donating to support the ongoing development.

Support Us