Podcast Episodes

Back to Search
Hotwiring Apple's Neural Engine
Hotwiring Apple's Neural Engine

Apple’s Neural Engine is one of the most powerful, and least accessible, AI accelerators in consumer hardware. In this episode of Neural Intel, we di…

3 weeks, 4 days ago

Short Long
View Episode
2026 LLM Inference Deep Dive: Solving the Memory Bandwidth & Interconnect Bottleneck | Neural Intel
2026 LLM Inference Deep Dive: Solving the Memory Bandwidth & Interconnect Bottleneck | Neural Intel

"Tokens per second screenshots are not architecture."

If you’re building sovereign AI systems, you need to understand why decode is memory-bandwidth-…

1 month, 1 week ago

Short Long
View Episode
Engineering Persistence: How MLX-Engine v1.8.5 Solves the KV Cache Rewind Problem
Engineering Persistence: How MLX-Engine v1.8.5 Solves the KV Cache Rewind Problem

Welcome back to Neural Intel. Today, we are going deep into the weeds of mlx-engine v1.8.5, the MIT-licensed inference backend for LM Studio.Neural S…

1 month, 1 week ago

Short Long
View Episode
Claude Fable 5 Isn’t Just a Better Model: It’s a New AI Runtime
Claude Fable 5 Isn’t Just a Better Model: It’s a New AI Runtime

Claude Fable 5 looks like a model launch on the surface. But underneath, the more interesting story is about runtime design: long-context workflows, …

1 month, 3 weeks ago

Short Long
View Episode
The EML Operator: One Primitive to Rule All Mathematics
The EML Operator: One Primitive to Rule All Mathematics

In this episode of Neural Intel, we perform a technical extraction of the paper "All elementary functions from a single operator". We discuss the sys…

2 months, 2 weeks ago

Short Long
View Episode
OpenAI MRC, SRv6, and the Architecture of Frontier AI Supercomputers
OpenAI MRC, SRv6, and the Architecture of Frontier AI Supercomputers

In this episode of the Neural Intel podcast, we go under the hood of OpenAI’s latest networking contribution to the Open Compute Project (OCP). We an…

2 months, 3 weeks ago

Short Long
View Episode
Inside the Machine: Training GPT-5, the Memory Wall, and the Math of MoE
Inside the Machine: Training GPT-5, the Memory Wall, and the Math of MoE

How are the world's most advanced models-GPT-5, Claude, and Gemini-actually trained and served at scale? In this deep dive, we move to the blackboard…

3 months ago

Short Long
View Episode
DeepSeek-V4: The Million-Token Efficiency Leap | Open Source SOTA
DeepSeek-V4: The Million-Token Efficiency Leap | Open Source SOTA

DeepSeek-AI has just dropped the DeepSeek-V4 series, featuring a massive 1.6T parameter MoE model that natively supports a one-million-token context …

3 months ago

Short Long
View Episode
Breaking the Quadratic Bottleneck with DeepSeek-V4’s Hybrid Attention
Breaking the Quadratic Bottleneck with DeepSeek-V4’s Hybrid Attention

Welcome back to the Neural Intel podcast. In this episode, we conduct a deep Neural Signal Check on the DeepSeek-V4 series to understand the architec…

3 months ago

Short Long
View Episode
Claude Desktop’s Silent Sandbox Bypass: The Undocumented Browser Bridge
Claude Desktop’s Silent Sandbox Bypass: The Undocumented Browser Bridge

Anthropic has been caught silently installing a Native Messaging manifest across seven different Chromium-based browsers, even those not present on y…

3 months, 1 week ago

Short Long
View Episode

Love PodBriefly?

If you like Podbriefly.com, please consider donating to support the ongoing development.

Support Us