Podcast Episodes
Back to SearchEP143: Context Engineering — Keep Long AI Requests Useful and Affordable
A practical guide to context engineering for AI APIs: select relevant information, manage long inputs, preserve important state, control token costs,…
1 month, 1 week ago
EP339: AI API Gateway Sampling Governance - Keep the Traces That Explain Reality
Govern AI API observability sampling across head and tail decisions, rare failures, retries, queues, tenants, redaction, incident windows, fairness, …
1 month, 1 week ago
EP338: AI API Gateway Tokenizer Governance - Keep Counting, Limits, and Billing Aligned
Govern AI API tokenizer and chat-template versions across admission, tools, multimodal units, context limits, routing, truncation, provider usage, re…
1 month, 1 week ago
EP337: AI API Gateway Encryption Contexts - Bind Data to the Right Key, Purpose, and Region
Govern AI API encryption contexts across tenant, purpose, artifact version, region, key rotation, queues, caches, providers, backups, deletion, and l…
1 month, 1 week ago
EP336: AI API Gateway Context Propagation - Carry Trust and Deadlines Across Every Hop
Propagate AI API request identity, tracing, deadlines, cancellation, priority, policy, residency, and causality safely across queues, workers, provid…
1 month, 1 week ago
EP335: AI API Gateway Response Delivery - Confirm What the Client Actually Received
Design reliable AI API response delivery with explicit lifecycle states, durable persistence, streaming final events, resumable retrieval, webhook se…
1 month, 1 week ago
EP334: AI API Gateway Capability Declarations - Route by What Providers Can Prove
Make AI provider capabilities verifiable for routing with typed, versioned facts covering limits, feature combinations, regions, account scope, probe…
1 month, 1 week ago
EP333: AI API Gateway Prompt Assembly - Preserve Source, Priority, and Boundaries
Govern AI prompt assembly with typed fragments, explicit precedence, provenance, tenant scope, retrieval and tool boundaries, multimodal identity, pr…
1 month, 1 week ago
EP332: AI API Gateway Model Alias Governance - Resolve Friendly Names Without Losing Control
Govern AI model aliases across version resolution, capabilities, rollout, routing, caching, billing, regional availability, fallback, rollback, and a…
1 month, 1 week ago
EP331: AI API Gateway Request Compression - Move Fewer Bytes Without Changing Meaning
Compress AI API gateway requests safely with explicit encodings, bounded expansion, streaming decompression, versioned signatures, strict parsing, ar…
1 month, 1 week ago