Episode Details
Back to Episodes
S13 Bonus: The Enterprise AI Chip War: Rethinking LLM Silicon & Inference with Vasanth Mohan, Director of Product at SambaNova
Season 13
Published 1 week, 4 days ago
Description
Today, we welcome a special guest to the podcast, Vasanth Mohan, Head of Developer Relations and Product Marking at SambaNova.
SambaNova is transforming AI with efficiency, security and sovereignty, driven by their relentless pursuit of intelligence, and running the largest models by maximizing dataflow efficiency. Vasanth and I dig into some fun - and heavy - topics at the backbone of where the industry is going with AI Inference, specifically at the hardware layer.
Questions:
- Tell me and the audience a bit about you and SambaNova.
- The Shift to Autonomy: "When we shift from a user generating a single chat response to an autonomous coding agent executing dozens of background tool calls, loop checks, and file rewrites, how does the underlying inference profile change, and what metric breaks first?
- Latency Budgets: "Token-per-second throughput used to be a nice-to-have metric for human readability, but for multi-agent workflows running in parallel, low latency is critical to prevent system timeouts. How are developers designing prompt structures and context windows to prevent compounding latencies during multi-agent orchestration?
- Hardware Heterogeneity: "Hyperscalers historically standardized on monolithic hardware, but neoclouds are increasingly mixing high-memory GPUs, specialized inference ASICs, and custom interconnects. How do you decide the optimal hardware mix inside a rack when your customer demand fluctuates between massive long-context reasoning models and rapid edge-like decoding?
- Power and Rack Density: "With high-end inference accelerators drawing massive power per node, how are neocloud data centers re-engineering liquid cooling, rack layouts, and power distribution specifically to maximize inference density rather than training throughput?
- Decoupling Prefill and Decode: "Disaggregated inference splits compute-heavy prefill operations from memory-bandwidth-heavy decode operations onto distinct hardware pools. What are the biggest real-world friction points when deploying this in production—especially around KV-cache transfer overhead and inter-node networking?
- Dynamic Orchestration: "In a disaggregated setup, workload spikes in long-context coding prompts can saturate the prefill pool while leaving decode nodes underutilized. How are leading engineering teams building smart routing layers and dynamic auto-scaling engines to balance compute across compute-bound and memory-bound hardware?
- The Economic Tipping Point: "Enterprise adoption is oscillating between proprietary API-driven models and self-hosted open-weights models. At what scale—measured in inference volume, latency SLAs, or data privacy constraints—does it become economically imperative for a company to transition off closed APIs onto self-hosted open infrastructure?
Links
https://www.linkedin.com/in/v-mohan/
Current Sponsors:
Checkout our Stacklist! https://stacks.codestory.co/
Hosted by Noah Labhart | Technical Founder & Startup Mentor.
Our Sponsors:
* Check out Granola and use my code g