Podcast Episodes
Back to SearchUltraQuant: 4-bit KV Caching for Context-Heavy Agents
## Episode Summary In this episode, we cover: - **UltraQuant: 4-bit KV Caching for Context-Heavy Agents** (arXiv) - **Fractional Verkle Trees: A Hype…
3 months, 2 weeks ago
Low-Energy Reduced RISC-V Instruction Subset Processor for Tsetlin Machine Inference at the Edge
## Episode Summary In this episode, we cover: - **Low-Energy Reduced RISC-V Instruction Subset Processor for Tsetlin Machine Inference at the Edge** …
3 months, 2 weeks ago
AIA: A Customized Multi-core RISC-V SoC for Discrete Sampling Workloads in 16 nm
## Episode Summary In this episode, we cover: - **AIA: A Customized Multi-core RISC-V SoC for Discrete Sampling Workloads in 16 nm** (arXiv) - **AIA:…
3 months, 2 weeks ago
SMEPilot: Characterizing and Optimizing LLM Inference with Scalable Matrix Extensions
## Episode Summary In this episode, we cover: - **SMEPilot: Characterizing and Optimizing LLM Inference with Scalable Matrix Extensions** (arXiv) - *…
3 months, 3 weeks ago
A Spatio-Temporal Expert Prefetching Framework for Efficient MoE-based LLM Inference
## Episode Summary In this episode, we cover: - **A Spatio-Temporal Expert Prefetching Framework for Efficient MoE-based LLM Inference** (arXiv) - **…
3 months, 3 weeks ago
Tiara: A Programmable Line-Rate ISA for Remote Memory Access
## Episode Summary In this episode, we cover: - **Tiara: A Programmable Line-Rate ISA for Remote Memory Access** (arXiv) - **HierSVA: A Data Synthesi…
3 months, 3 weeks ago
Isolation-aware Scheduling Framework for DNN-based End-to-End Autonomous Driving System on Tile-based Accelerators
## Episode Summary In this episode, we cover: - **Isolation-aware Scheduling Framework for DNN-based End-to-End Autonomous Driving System on Tile-bas…
3 months, 3 weeks ago
Energy-Efficient On-Device RAG on a Mobile NPU: System Design and Benchmark on Snapdragon X Elite
## Episode Summary In this episode, we cover: - **Energy-Efficient On-Device RAG on a Mobile NPU: System Design and Benchmark on Snapdragon X Elite**…
3 months, 3 weeks ago
SupraSNN: Exploiting Synapse-Level Parallelism in Spiking Neural Network Accelerators through Co-Optimized Mapping and Scheduling
## Episode Summary In this episode, we cover: - **SupraSNN: Exploiting Synapse-Level Parallelism in Spiking Neural Network Accelerators through Co-Op…
3 months, 4 weeks ago
NextSilicon to Productize Arbel RISC-V Core Into 64-Core Enterprise Processor for AI and HPC - Business Wire
## Episode Summary In this episode, we cover: - **NextSilicon to Productize Arbel RISC-V Core Into 64-Core Enterprise Processor for AI and HPC - Busi…
3 months, 4 weeks ago