Episode Details
Back to EpisodesPipeline-Native Transformers: Co-Designing Model Architecture and CPU Inference for Bandwidth-Efficient Autoregressive Decode
Published 1 month, 2 weeks ago
Description
## Episode Summary
In this episode, we cover:
- **Pipeline-Native Transformers: Co-Designing Model Architecture and CPU Inference for Bandwidth-Efficient Autoregressive Decode** (arXiv)
- **Elastic KV Cache for LLM Serving:A Working Reclamation Mechanism, and Why Chunked Prefill Already Closes the Gap** (arXiv)
- **Design of a low-power RISC-V based intelligent endoscopy detection processor: EndoRISC-V - Nature** (google_riscv)
- **Prompt to tape out: Autonomous AI agent builds 1.5 GHz RISC-V CPU - Adafruit** (google_riscv)
- **Armaan Gomes and Friends Deliver Playable Doom on a Whole New Platform: A Custom RISC-V CPU - Hackster.io** (google_riscv)
---
*Sponsored by Ada, Ago Consulting, and Zen Semiconductor*