Podcast Episodes

Back to Search
UltraQuant: 4-bit KV Caching for Context-Heavy Agents

## Episode Summary In this episode, we cover: - **UltraQuant: 4-bit KV Caching for Context-Heavy Agents** (arXiv) - **Fractional Verkle Trees: A Hype…

3 months, 2 weeks ago

Short Long
View Episode
Low-Energy Reduced RISC-V Instruction Subset Processor for Tsetlin Machine Inference at the Edge

## Episode Summary In this episode, we cover: - **Low-Energy Reduced RISC-V Instruction Subset Processor for Tsetlin Machine Inference at the Edge** …

3 months, 2 weeks ago

Short Long
View Episode
AIA: A Customized Multi-core RISC-V SoC for Discrete Sampling Workloads in 16 nm

## Episode Summary In this episode, we cover: - **AIA: A Customized Multi-core RISC-V SoC for Discrete Sampling Workloads in 16 nm** (arXiv) - **AIA:…

3 months, 2 weeks ago

Short Long
View Episode
SMEPilot: Characterizing and Optimizing LLM Inference with Scalable Matrix Extensions

## Episode Summary In this episode, we cover: - **SMEPilot: Characterizing and Optimizing LLM Inference with Scalable Matrix Extensions** (arXiv) - *…

3 months, 3 weeks ago

Short Long
View Episode
A Spatio-Temporal Expert Prefetching Framework for Efficient MoE-based LLM Inference

## Episode Summary In this episode, we cover: - **A Spatio-Temporal Expert Prefetching Framework for Efficient MoE-based LLM Inference** (arXiv) - **…

3 months, 3 weeks ago

Short Long
View Episode
Tiara: A Programmable Line-Rate ISA for Remote Memory Access

## Episode Summary In this episode, we cover: - **Tiara: A Programmable Line-Rate ISA for Remote Memory Access** (arXiv) - **HierSVA: A Data Synthesi…

3 months, 3 weeks ago

Short Long
View Episode
Isolation-aware Scheduling Framework for DNN-based End-to-End Autonomous Driving System on Tile-based Accelerators

## Episode Summary In this episode, we cover: - **Isolation-aware Scheduling Framework for DNN-based End-to-End Autonomous Driving System on Tile-bas…

3 months, 3 weeks ago

Short Long
View Episode
Energy-Efficient On-Device RAG on a Mobile NPU: System Design and Benchmark on Snapdragon X Elite

## Episode Summary In this episode, we cover: - **Energy-Efficient On-Device RAG on a Mobile NPU: System Design and Benchmark on Snapdragon X Elite**…

3 months, 3 weeks ago

Short Long
View Episode
SupraSNN: Exploiting Synapse-Level Parallelism in Spiking Neural Network Accelerators through Co-Optimized Mapping and Scheduling

## Episode Summary In this episode, we cover: - **SupraSNN: Exploiting Synapse-Level Parallelism in Spiking Neural Network Accelerators through Co-Op…

3 months, 4 weeks ago

Short Long
View Episode
NextSilicon to Productize Arbel RISC-V Core Into 64-Core Enterprise Processor for AI and HPC - Business Wire

## Episode Summary In this episode, we cover: - **NextSilicon to Productize Arbel RISC-V Core Into 64-Core Enterprise Processor for AI and HPC - Busi…

3 months, 4 weeks ago

Short Long
View Episode

Love PodBriefly?

If you like Podbriefly.com, please consider donating to support the ongoing development.

Support Us