Beyond Fast Contractions: Attenuation and Recovery of Matrix-Engine Speedups in High-Order Finite Elements

2026-08-10Distributed, Parallel, and Cluster Computing

Distributed, Parallel, and Cluster Computing
AI summary

The authors study why newer processors with powerful matrix engines don't always make scientific computing faster in practice. They focus on a key part of a simulation code running on specific Arm CPUs and find that while the matrix engine is much faster at raw calculations, the overall speedup is much smaller. They identify several bottlenecks, like data movement and synchronization, that limit gains. By changing how data is organized and processed, the authors improve speed but show that full benefits need redesigning the entire calculation process, not just speeding up one part.

matrix enginesSIMD (Single Instruction Multiple Data)SVE (Scalable Vector Extension)tensor contractionsscientific computingdata movementsynchronizationoperator optimizationArm LX2 CPUperformance bottlenecks
Authors
Yinuo Wang, Lin Gan, Tianqi Mao, Zeyu Song, Wubing Wan, Jiayu Fu, Zekun Yin, Yuyang Jin, Xiaohui Duan, Wei Xue, Guangwen Yang
Abstract
Modern processors increasingly provide matrix engines whose peak arithmetic throughput greatly exceeds conventional SIMD, but scientific applications rarely realize this advantage end to end. We examine this gap in SPECFEM3D's dominant stiffness operator on the Arm LX2 CPUs that power the flagship Lineshine supercomputer. Against a matched, high-performance SVE baseline on the same cores, SME's $4\times$ single-precision peak advantage falls to $2.2\times$ for isolated tensor contractions and $1.1\times$ for the complete operator. Our factorized diagnostic attributes the loss to pointwise computation, indirect field movement and synchronization, and irregular coefficient delivery. Explicit SIMD mitigates pointwise work, raising the full-operator speedup to $1.3\times$. Field-layout changes mitigate indirect movement and synchronization, while vector-blocked coefficient streaming reduces irregular-access costs; together they raise speedup to $1.6\times$ at high order. A contraction-free control bounds further contraction-only gains at $1.11$--$1.32\times$. Realizing matrix-engine performance therefore requires co-designing the entire operator path, not merely replacing its contraction kernel.