Podcast Episodes
Back to SearchThe key CPU architecture you've never heard of: RISC-V - redsharknews.com
## Episode Summary In this episode, we cover: - **The key CPU architecture you've never heard of: RISC-V - redsharknews.com** (google_arch) - **Check…
1 month ago
SchedBlame: Who Ran While You Waited? Culprit-Attributed CPU Contention for Containers on Stock Kernels
## Episode Summary In this episode, we cover: - **SchedBlame: Who Ran While You Waited? Culprit-Attributed CPU Contention for Containers on Stock Ker…
1 month, 1 week ago
Performance Characterization of SPEC CPU 2026 on AMD EPYC 9755 Processor
## Episode Summary In this episode, we cover: - **Performance Characterization of SPEC CPU 2026 on AMD EPYC 9755 Processor** (arXiv) - **Hardware Acc…
1 month, 1 week ago
CHIPSMORE: Compute-in-Interconnect and -Memory Chiplets for Multi-Mode Multi-Request LLM Inference Acceleration
## Episode Summary In this episode, we cover: - **CHIPSMORE: Compute-in-Interconnect and -Memory Chiplets for Multi-Mode Multi-Request LLM Inference …
1 month, 1 week ago
Great Expectations: Benchmarking the Real-World Performance of RVV 1.0 in HPC
## Episode Summary In this episode, we cover: - **Great Expectations: Benchmarking the Real-World Performance of RVV 1.0 in HPC** (arXiv) - **AI Hard…
1 month, 1 week ago
NOVA: Technology-Architecture Co-Design of Near-Memory Processing for Attention-SSM-MoE Hybrid LLM Inference
## Episode Summary In this episode, we cover: - **NOVA: Technology-Architecture Co-Design of Near-Memory Processing for Attention-SSM-MoE Hybrid LLM …
1 month, 1 week ago
SILK: Closing the Time-of-Check-to-Time-of-Use Gap in RoT-Protected AI Systems
## Episode Summary In this episode, we cover: - **SILK: Closing the Time-of-Check-to-Time-of-Use Gap in RoT-Protected AI Systems** (arXiv) - **Simthe…
1 month, 1 week ago
PRO-RAN: Processor-Level Characterization of Open RAN Centralized and Distributed Units
## Episode Summary In this episode, we cover: - **PRO-RAN: Processor-Level Characterization of Open RAN Centralized and Distributed Units** (arXiv) -…
1 month, 1 week ago
FLINT: Efficiently Leveraging High Bandwidth Flash for Capacity-Scalable LLM Inference Acceleration
## Episode Summary In this episode, we cover: - **FLINT: Efficiently Leveraging High Bandwidth Flash for Capacity-Scalable LLM Inference Acceleration…
1 month, 1 week ago
Pipeline-Native Transformers: Co-Designing Model Architecture and CPU Inference for Bandwidth-Efficient Autoregressive Decode
## Episode Summary In this episode, we cover: - **Pipeline-Native Transformers: Co-Designing Model Architecture and CPU Inference for Bandwidth-Effic…
1 month, 2 weeks ago