Podcast Episodes
Back to SearchArchitecting the Next Generation of Asynchronous, Distributed GPUs for the AI Era
## Episode Summary In this episode, we cover: - **Architecting the Next Generation of Asynchronous, Distributed GPUs for the AI Era** (arXiv) - **Tre…
1 month, 2 weeks ago
SPICE: Speculative Prefetching with Low-Rank Expert Surrogates and Heterogeneous Orchestration for MoE Inference Acceleration
## Episode Summary In this episode, we cover: - **SPICE: Speculative Prefetching with Low-Rank Expert Surrogates and Heterogeneous Orchestration for …
1 month, 2 weeks ago
Formal Performance and Compile Time Guarantees for Compiler Optimization Heuristics
## Episode Summary In this episode, we cover: - **Formal Performance and Compile Time Guarantees for Compiler Optimization Heuristics** (arXiv) - **B…
1 month, 2 weeks ago
From Retrieved Context to Runtime Control: Adaptive Compression for Edge-based RAG
## Episode Summary In this episode, we cover: - **From Retrieved Context to Runtime Control: Adaptive Compression for Edge-based RAG** (arXiv) - **DT…
1 month, 2 weeks ago
A Thread-Register Decoupled GPU Execution Model for Efficient Tensor Computation
## Episode Summary In this episode, we cover: - **A Thread-Register Decoupled GPU Execution Model for Efficient Tensor Computation** (arXiv) - **Perf…
1 month, 2 weeks ago
FPGA Lifecycle Management for RISC-V Systems
## Episode Summary In this episode, we cover: - **FPGA Lifecycle Management for RISC-V Systems** (arXiv) - **CTTE: An Open Dual-Protocol RISC-V Trace…
1 month, 2 weeks ago
ETHEREAL: A 25.6-$μ$s/inf. Low-latency Event-driven Graph-neural-network Processor for High-resolution Vision at the Edge
## Episode Summary In this episode, we cover: - **ETHEREAL: A 25.6-$μ$s/inf. Low-latency Event-driven Graph-neural-network Processor for High-resolut…
1 month, 3 weeks ago
Testing the EPYC Conjecture on Real Hardware: MoA-Guided Dense Matrix Multiplication on NCSA Delta (AMD EPYC 7763 Milan)
## Episode Summary In this episode, we cover: - **Testing the EPYC Conjecture on Real Hardware: MoA-Guided Dense Matrix Multiplication on NCSA Delta …
1 month, 3 weeks ago
SLAC: Access-Driven CPU-to-GPU Side-channel Attacks via System-Level Cache on Apple Silicon
## Episode Summary In this episode, we cover: - **SLAC: Access-Driven CPU-to-GPU Side-channel Attacks via System-Level Cache on Apple Silicon** (arXi…
1 month, 3 weeks ago
What Actually Serializes GPU LZ77 Decode: Three Decoders, Three Mechanisms, and an Encode-Time Lever That Removes the Last One
## Episode Summary In this episode, we cover: - **What Actually Serializes GPU LZ77 Decode: Three Decoders, Three Mechanisms, and an Encode-Time Leve…
1 month, 3 weeks ago