Episode Details

Back to Episodes

STEEL: Sparsity-Aware Fused Attention for Energy-Efficient Long-Sequence Inference on AMD's XDNA NPU

Published 2 months, 3 weeks ago
Description
## Episode Summary In this episode, we cover: - **STEEL: Sparsity-Aware Fused Attention for Energy-Efficient Long-Sequence Inference on AMD's XDNA NPU** (arXiv) - **StreamDQ: Near-Memory Weight DeQuantization in Custom HBM for Scalable AI Inference Acceleration** (arXiv) - **Security Researchers Find Current RISC-V CPU Implementations Coming Up Short - Phoronix** (google_riscv) - **SiFive Upgrades P500 RISC-V IP Cores - embedded.com** (google_riscv) - **Vividnode Mobile AI Packs RISC-V Processor and 60 TOPS AI Engine - LinuxGizmos.com** (google_riscv) --- *Sponsored by LimitLess AI*
Listen Now

Love PodBriefly?

If you like Podbriefly.com, please consider donating to support the ongoing development.

Support Us