Podcast Episodes
Back to SearchCacheFit: The Rising Cost of Cache Replacement and Constraint-Driven Redesign of Last-Level Cache Replacement Policies
## Episode Summary In this episode, we cover: - **CacheFit: The Rising Cost of Cache Replacement and Constraint-Driven Redesign of Last-Level Cache R…
7 hours ago
CHiRP: Control-Flow History Reuse Prediction
## Episode Summary In this episode, we cover: - **CHiRP: Control-Flow History Reuse Prediction** (arXiv) - **SVRF: Efficient Register Storage for Lon…
1 day, 7 hours ago
Trail: Scalable and Low-Cost Temporal TLB Prefetching via Page-Table-Embedded Deltas
## Episode Summary In this episode, we cover: - **Trail: Scalable and Low-Cost Temporal TLB Prefetching via Page-Table-Embedded Deltas** (arXiv) - **…
2 days, 7 hours ago
Characterizing High Bandwidth Flash for LLM Serving
## Episode Summary In this episode, we cover: - **Characterizing High Bandwidth Flash for LLM Serving** (arXiv) - **Request Order Matters: Cache-Hist…
3 days, 7 hours ago
ServeTwin: A Benchmark-Validated Simulator for Distributed LLM Architecture Exploration
## Episode Summary In this episode, we cover: - **ServeTwin: A Benchmark-Validated Simulator for Distributed LLM Architecture Exploration** (arXiv) -…
4 days, 7 hours ago
Where Activation Sparsity and KV-Cache Sparsity Cross in LLM Decoding
## Episode Summary In this episode, we cover: - **Where Activation Sparsity and KV-Cache Sparsity Cross in LLM Decoding** (arXiv) - **U-Sonic: An Ope…
5 days, 7 hours ago
Efficient Linkage-Based Compartmentalization on CHERI
## Episode Summary In this episode, we cover: - **Efficient Linkage-Based Compartmentalization on CHERI** (arXiv) - **SPIMOE: Exploiting Hybrid Spars…
6 days, 7 hours ago
Characterizing High Bandwidth Flash for LLM Serving
## Episode Summary In this episode, we cover: - **Characterizing High Bandwidth Flash for LLM Serving** (arXiv) - **STELLA: A 16nm Spatio-Temporal El…
1 week ago
Cache-Aware Conv3D Lowering Across Embedded World-Model Decoders
## Episode Summary In this episode, we cover: - **Cache-Aware Conv3D Lowering Across Embedded World-Model Decoders** (arXiv) - **DPS: Dual-Mode Preci…
1 week, 1 day ago
Bandwidth, Not FLOPS: FFT Kernels, Matrix Units and SAR Imaging on Apple M6
## Episode Summary In this episode, we cover: - **Bandwidth, Not FLOPS: FFT Kernels, Matrix Units and SAR Imaging on Apple M6** (arXiv) - **Improving…
1 week, 2 days ago