Episode Details

Back to Episodes
What Precision Is Your Model Actually Running At?

What Precision Is Your Model Actually Running At?

Episode 5004 Published 1 month ago
Description
When you call an API for GPT-4o or Llama 3 70B, are you actually getting the full-precision model you think you are? This episode pulls back the curtain on inference infrastructure, examining how commercial providers like Together AI and Fireworks quantize open-weight models by default, whether closed-source vendors like OpenAI and Anthropic optimize their own models in undisclosed ways, and what happens when your request routes through aggregators like OpenRouter. We explore the economic pressure to quantize, the documented vs. undocumented reality of production inference, and what it means for developers who assume the weights are untouched. Episode #692729 — open it directly at myweirdprompts.com/692729
Listen Now

Love PodBriefly?

If you like Podbriefly.com, please consider donating to support the ongoing development.

Support Us