Podcast Episodes

Back to Search
CacheFit: The Rising Cost of Cache Replacement and Constraint-Driven Redesign of Last-Level Cache Replacement Policies

## Episode Summary In this episode, we cover: - **CacheFit: The Rising Cost of Cache Replacement and Constraint-Driven Redesign of Last-Level Cache R…

7 hours ago

Short Long
View Episode
CHiRP: Control-Flow History Reuse Prediction

## Episode Summary In this episode, we cover: - **CHiRP: Control-Flow History Reuse Prediction** (arXiv) - **SVRF: Efficient Register Storage for Lon…

1 day, 7 hours ago

Short Long
View Episode
Trail: Scalable and Low-Cost Temporal TLB Prefetching via Page-Table-Embedded Deltas

## Episode Summary In this episode, we cover: - **Trail: Scalable and Low-Cost Temporal TLB Prefetching via Page-Table-Embedded Deltas** (arXiv) - **…

2 days, 7 hours ago

Short Long
View Episode
Characterizing High Bandwidth Flash for LLM Serving

## Episode Summary In this episode, we cover: - **Characterizing High Bandwidth Flash for LLM Serving** (arXiv) - **Request Order Matters: Cache-Hist…

3 days, 7 hours ago

Short Long
View Episode
ServeTwin: A Benchmark-Validated Simulator for Distributed LLM Architecture Exploration

## Episode Summary In this episode, we cover: - **ServeTwin: A Benchmark-Validated Simulator for Distributed LLM Architecture Exploration** (arXiv) -…

4 days, 7 hours ago

Short Long
View Episode
Where Activation Sparsity and KV-Cache Sparsity Cross in LLM Decoding

## Episode Summary In this episode, we cover: - **Where Activation Sparsity and KV-Cache Sparsity Cross in LLM Decoding** (arXiv) - **U-Sonic: An Ope…

5 days, 7 hours ago

Short Long
View Episode
Efficient Linkage-Based Compartmentalization on CHERI

## Episode Summary In this episode, we cover: - **Efficient Linkage-Based Compartmentalization on CHERI** (arXiv) - **SPIMOE: Exploiting Hybrid Spars…

6 days, 7 hours ago

Short Long
View Episode
Characterizing High Bandwidth Flash for LLM Serving

## Episode Summary In this episode, we cover: - **Characterizing High Bandwidth Flash for LLM Serving** (arXiv) - **STELLA: A 16nm Spatio-Temporal El…

1 week ago

Short Long
View Episode
Cache-Aware Conv3D Lowering Across Embedded World-Model Decoders

## Episode Summary In this episode, we cover: - **Cache-Aware Conv3D Lowering Across Embedded World-Model Decoders** (arXiv) - **DPS: Dual-Mode Preci…

1 week, 1 day ago

Short Long
View Episode
Bandwidth, Not FLOPS: FFT Kernels, Matrix Units and SAR Imaging on Apple M6

## Episode Summary In this episode, we cover: - **Bandwidth, Not FLOPS: FFT Kernels, Matrix Units and SAR Imaging on Apple M6** (arXiv) - **Improving…

1 week, 2 days ago

Short Long
View Episode

Love PodBriefly?

If you like Podbriefly.com, please consider donating to support the ongoing development.

Support Us