Episode Details

Back to Episodes
DeepSeek's Point Release That Isn't

DeepSeek's Point Release That Isn't

Episode 5188 Published 1 month ago
Description
DeepSeek V4.1 Flash looks like a minor revision. It isn't. The layer structure is new, the KV cache was rebuilt from scratch, and nearly 200 billion parameters aren't transformer weights at all — they're lookup tables. This episode walks through the intermediate layer most AI commentary skips: what layers actually do, how the encoder/decoder split interacts with them, and where the KV cache sits in that picture. Along the way: why the cache shrank 437x since V1, what "You Only Cache Once" means in practice, and why a model that's cheaper to read than to write is exactly the shape agent workloads need. Episode #923768 — open it directly at myweirdprompts.com/923768
Listen Now

Love PodBriefly?

If you like Podbriefly.com, please consider donating to support the ongoing development.

Support Us