Podcast Episodes
Back to SearchQuMA: Researchers Develop Quantum Microarchitecture that "Bridges the Gap" in Processor System Stacks - IEEE Computer Society
## Episode Summary In this episode, we cover: - **QuMA: Researchers Develop Quantum Microarchitecture that "Bridges the Gap" in Processor System Stac…
1 week, 3 days ago
The KV Cache Is the New Memory Wall
## Episode Summary In this episode, we cover: - **The KV Cache Is the New Memory Wall** (arXiv) - **PipeDRAM: A Data-Transposition-Free In-DRAM Archi…
1 week, 4 days ago
HeteroReason: Heterogeneous FPGA-GPU Acceleration for Disaggregated Speculative Reasoning
## Episode Summary In this episode, we cover: - **HeteroReason: Heterogeneous FPGA-GPU Acceleration for Disaggregated Speculative Reasoning** (arXiv)…
1 week, 5 days ago
Exploiting Decompression Latency for Covert Channels in Inter-Line-Compressed LLCs
## Episode Summary In this episode, we cover: - **Exploiting Decompression Latency for Covert Channels in Inter-Line-Compressed LLCs** (arXiv) - **Pa…
1 week, 6 days ago
Bridging LLM Serving and CXL-SSDs with Chunk-Aware KV Cache Management
## Episode Summary In this episode, we cover: - **Bridging LLM Serving and CXL-SSDs with Chunk-Aware KV Cache Management** (arXiv) - **HBF-Sim: An Ex…
2 weeks ago
PrefixBench-H100: Characterizing Prefix Reuse and Time-to-First-Token in H100 LLM Serving
## Episode Summary In this episode, we cover: - **PrefixBench-H100: Characterizing Prefix Reuse and Time-to-First-Token in H100 LLM Serving** (arXiv)…
2 weeks, 6 days ago
RISC-V and machine learning: a survey
## Episode Summary In this episode, we cover: - **RISC-V and machine learning: a survey** (arXiv) - **MeshKV: A Network-on-Chip KV Cache Fabric for S…
3 weeks ago
Quantifying the Effect of HCLs on a Fixed-Microarchitecture MXFP4 Accelerator
## Episode Summary In this episode, we cover: - **Quantifying the Effect of HCLs on a Fixed-Microarchitecture MXFP4 Accelerator** (arXiv) - **HBFlex:…
3 weeks, 1 day ago
The World Model Hardware Accelerator
## Episode Summary In this episode, we cover: - **The World Model Hardware Accelerator** (arXiv) - **A 25-$μ$s/inf Event-driven Graph Neural Network …
3 weeks, 2 days ago
BOOST: Concurrent Access to Host Memory and HBM to Accelerate LLM Inference
## Episode Summary In this episode, we cover: - **BOOST: Concurrent Access to Host Memory and HBM to Accelerate LLM Inference** (arXiv) - **Arborist:…
3 weeks, 3 days ago