Episode Details
Back to Episodes
The Neural Deep Dive 2026-07-15: Cheating the Hardware Limit
Published 1 month, 2 weeks ago
Description
Break the link between model size and compute cost. We dive into the "Sparse Mixture-of-Experts" revolution, exploring how the Llama 4 Scout architecture uses dynamic routing and 4-bit quantization to bring trillion-parameter intelligence to local hardware. From "routing collapse" to the promise of decentralized AI, discover why smarter routing is replacing bigger models.