Episode Details
Back to EpisodesNOVA: Technology-Architecture Co-Design of Near-Memory Processing for Attention-SSM-MoE Hybrid LLM Inference
Published 1 month, 1 week ago
Description
## Episode Summary
In this episode, we cover:
- **NOVA: Technology-Architecture Co-Design of Near-Memory Processing for Attention-SSM-MoE Hybrid LLM Inference** (arXiv)
- **Performance Foundations of Parallel & Distributed Reasoning Language Models** (arXiv)
- **Nvidia Vera CPU Architecture: Max single-threaded CPU at scale for agents - GamesBeat** (google_arch)
- **NextSilicon to Productize Arbel RISC-V Core Into 64-Core Enterprise Processor for AI and HPC - Business Wire** (google_riscv)
- **NextSilicon to Productize Arbel RISC-V Core into 64-Core Enterprise Processor for AI and HPC - HPCwire** (google_riscv)
---
*Sponsored by Ada, Ago Consulting, and Zen Semiconductor*