Hybrid attention models run faster and use less energy on edge NPUs

Empowering Hybrid Attention Models on NPUs

Distributed, Parallel, and Cluster ComputingArtificial Intelligence

Summary

Large language models use a technique called hybrid attention to understand language efficiently, but running these on small edge devices with special chips called NPUs has challenges. The authors found that simply running these models on NPUs is slow and wasteful because of memory inefficiencies and mismatched hardware usage. They created a system called HA-NPU that reorganizes how the model's data and tasks flow through the hardware, making it much faster and more energy-efficient. This lets devices like smartphones or IoT gadgets run large language models more effectively without changing how the model itself works.

What this means in practice

  • For mobile app developers: Run advanced language models efficiently on smartphones using edge NPUs for faster response and lower energy use.$Commercial implications: Enables selling smarter AI-powered mobile apps requiring on-device large language model inference with better speed and battery life.
  • For embedded device engineers: Integrate large language model features in IoT devices by improving inference efficiency on edge NPUs without algorithm changes.

Authors

Yinyuan Zhang, Daliang Xu, Xiaolong Huang, Wangsong Yin, Yun Ma, Mengwei Xu, Gang Huang

Abstract

Hybrid attention models have emerged as a crucial architecture for Large Language Models (LLMs) (e.g., the Qwen3.5 and Kimi series). Their memory and computational efficiency make them highly attractive for on-device inference, forming a promising synergy with edge Neural Processing Units (NPUs). However, naive execution of these hybrid models on edge NPUs fails to deliver these benefits, often bottlenecking the prefill stage due to severe memory-system inefficiencies and architectural mismatches within the linear attention (LA) layers. We present HA-NPU, the first system to enable efficient hybrid attention LLM inference on edge NPUs without modifying the underlying algorithms. HA-NPU enhances execution efficiency by reorganizing the dataflow of the LA components across three levels: (1) At the core level, it partitions workloads by the head dimension and fuses dependent operators, eliminating cross-core global memory accesses; (2) At the operator level, it reorders execution to consume intermediate tensors immediately, drastically minimizing local-buffer pressure; (3) At the tensor level, it employs dataflow-aware layout planning to minimize transformation overhead between matrix and vector processing units. Compared to competitive baselines, HA-NPU achieves up to 35.95$\times$ LA kernel speedup and 36.14$\times$ energy reduction, delivering up to 2.03$\times$ faster end-to-end request latency. The source code will be made publicly available at https://github.com/yinyuanzhang/HA-NPU