Hybrid sparsity boosts reasoning model speed on memory processors
SPIMOE: Exploiting Hybrid Sparsity for Reasoning MoE Inference on Heterogeneous PIM Architectures
Hardware Architecture
Summary
Large AI models that use mixtures of experts (MoE) have slow parts when they try to remember and focus on important information, and also when they activate different expert parts unevenly. The authors created SPIMOE, a new system that uses a special way to handle sparsity—meaning only parts of the model are active at once—on processors that combine memory and computing. This system splits tasks between different memory-processing units and smartly manages the workload to speed up the model's thinking without losing accuracy. Their tests show SPIMOE is much faster than current GPUs and similar PIM-based methods.
What this means in practice
- •For ai model engineers: Speed up inference of large MoE models with uneven expert activity using hybrid sparsity on specialized memory processors.
- •For computer hardware architects: Design heterogeneous Processing-in-Memory architectures that efficiently share workloads for long-sequence reasoning models to improve utilization and throughput.
Authors
Rubing Yang, Cenlin Duan, Yingjie Qi, Xiaolin He, Xiao Ma, Jianlei Yang
Abstract
Long-reasoning Mixture-of-Experts (MoE) models expose two coupled inference bottlenecks: growing KV caches shift the critical path toward attention, while sparse expert activation causes load imbalance and low hardware utilization. Although Processing-in-Memory (PIM) offers a promising way to mitigate data movement overhead, existing PIM-based accelerators typically optimize attention or FFNs in isolation. We propose SPIMOE, the first co-design framework that exploits hybrid sparsity for efficient MoE inference on heterogeneous PIM architectures. SPIMOE combines adaptive expert routing with block-sparse attention and physical KV-cache eviction, and disaggregates attention and FFNs across SRAM-PIM and HBM-PIM. Static expert mapping and dynamic sub-batch scheduling further balance channel loads and overlap the two paths. Evaluations show that SPIMOE achieves up to $8.35\times$ end-to-end speedup over an NVIDIA A100 GPU and $3.33\times$ speedup in MoE FFN execution over PIMoE, while preserving reasoning accuracy comparable to full-attention baselines.