Episode Details

Back to Episodes

Pipeline-Native Transformers: Co-Designing Model Architecture and CPU Inference for Bandwidth-Efficient Autoregressive Decode

Published 1 month, 2 weeks ago
Description
## Episode Summary In this episode, we cover: - **Pipeline-Native Transformers: Co-Designing Model Architecture and CPU Inference for Bandwidth-Efficient Autoregressive Decode** (arXiv) - **Elastic KV Cache for LLM Serving:A Working Reclamation Mechanism, and Why Chunked Prefill Already Closes the Gap** (arXiv) - **Design of a low-power RISC-V based intelligent endoscopy detection processor: EndoRISC-V - Nature** (google_riscv) - **Prompt to tape out: Autonomous AI agent builds 1.5 GHz RISC-V CPU - Adafruit** (google_riscv) - **Armaan Gomes and Friends Deliver Playable Doom on a Whole New Platform: A Custom RISC-V CPU - Hackster.io** (google_riscv) --- *Sponsored by Ada, Ago Consulting, and Zen Semiconductor*
Listen Now

Love PodBriefly?

If you like Podbriefly.com, please consider donating to support the ongoing development.

Support Us