Papers for

computer hardware architects

Papers whose findings have a practical use for this group, as judged from the abstract. Open a paper to read what it means in practice.

Hybrid sparsity boosts reasoning model speed on memory processors

SPIMOE: Exploiting Hybrid Sparsity for Reasoning MoE Inference on Heterogeneous PIM Architectures

Abstract: Long-reasoning Mixture-of-Experts (MoE) models expose two coupled inference bottlenecks: growing KV caches shift the critical path toward attention, while sparse expert activation causes load imbalance and low hardware utilization. Although Processing-in-Memory (PIM) offers a promising way to mitigate data movement overhead, existing PIM-based accelerators typically optimize attention or FFNs in isolation. We propose SPIMOE, the first co-design framework that exploits hybrid sparsity for efficient MoE inference on heterogeneous PIM architectures. SPIMOE combines adaptive expert routing with block-sparse attention and physical KV-cache eviction, and disaggregates attention and FFNs across SRAM-PIM and HBM-PIM. Static expert mapping and dynamic sub-batch scheduling further balance channel loads and overlap the two paths. Evaluations show that SPIMOE achieves up to $8.35\times$ end-to-end speedup over an NVIDIA A100 GPU and $3.33\times$ speedup in MoE FFN execution over PIMoE, while preserving reasoning accuracy comparable to full-attention baselines.

Mon 28 SeptHardware Architecture
The gist
Large AI models that use mixtures of experts (MoE) have slow parts when they try to remember and focus on important information, and also when they activate different expert parts unevenly. The authors created SPIMOE, a new system that uses a special way to handle sparsity—meaning only parts of the model are active at once—on processors that combine memory and computing. This system splits tasks between different memory-processing units and smartly manages the workload to speed up the model's thinking without losing accuracy. Their tests show SPIMOE is much faster than current GPUs and similar PIM-based methods.
Open → 2609.34612v1