Episode Details

Back to Episodes

NOVA: Technology-Architecture Co-Design of Near-Memory Processing for Attention-SSM-MoE Hybrid LLM Inference

Published 1 month, 1 week ago
Description
## Episode Summary In this episode, we cover: - **NOVA: Technology-Architecture Co-Design of Near-Memory Processing for Attention-SSM-MoE Hybrid LLM Inference** (arXiv) - **Performance Foundations of Parallel & Distributed Reasoning Language Models** (arXiv) - **Nvidia Vera CPU Architecture: Max single-threaded CPU at scale for agents - GamesBeat** (google_arch) - **NextSilicon to Productize Arbel RISC-V Core Into 64-Core Enterprise Processor for AI and HPC - Business Wire** (google_riscv) - **NextSilicon to Productize Arbel RISC-V Core into 64-Core Enterprise Processor for AI and HPC - HPCwire** (google_riscv) --- *Sponsored by Ada, Ago Consulting, and Zen Semiconductor*
Listen Now

Love PodBriefly?

If you like Podbriefly.com, please consider donating to support the ongoing development.

Support Us