Episode Details
Back to EpisodesSTEEL: Sparsity-Aware Fused Attention for Energy-Efficient Long-Sequence Inference on AMD's XDNA NPU
Published 2 months, 3 weeks ago
Description
## Episode Summary
In this episode, we cover:
- **STEEL: Sparsity-Aware Fused Attention for Energy-Efficient Long-Sequence Inference on AMD's XDNA NPU** (arXiv)
- **StreamDQ: Near-Memory Weight DeQuantization in Custom HBM for Scalable AI Inference Acceleration** (arXiv)
- **Security Researchers Find Current RISC-V CPU Implementations Coming Up Short - Phoronix** (google_riscv)
- **SiFive Upgrades P500 RISC-V IP Cores - embedded.com** (google_riscv)
- **Vividnode Mobile AI Packs RISC-V Processor and 60 TOPS AI Engine - LinuxGizmos.com** (google_riscv)
---
*Sponsored by LimitLess AI*