Memory optimal transformer kernels show varied speed on different HPC hardware

Validating Memory-Optimal Transformer Kernels on Real Hardware: From Formal Derivation to Measured Performance Across Two HPC Clusters

Distributed, Parallel, and Cluster ComputingArtificial Intelligence

Summary

Transformer models, which power many AI systems, use a lot of computer memory and computing resources. The authors confirmed earlier math-based predictions about the best ways to run parts of these models using real high-performance computers. They fixed a slowdown problem on GPUs caused by how tasks were combined, showed that hardware layout drastically changes speed, and found surprising performance differences between C and Fortran on GPUs. Their approach helps adapt AI computations efficiently as hardware evolves, without redoing complex proofs.

What this means in practice

  • For high performance computing engineers: Optimize transformer model computations on specific HPC clusters by tailoring kernel implementations to hardware topology for improved performance.
  • For compiler and runtime developers: Use verified hardware-independent derivations as fixed blueprints and perform hardware-specific rewrites to generate efficient transformer kernels without re-deriving correctness.

Authors

Lenore M. Mullin, Gaetan Hains

Abstract

We validate memory-optimal cost functions for transformer kernels derived via the Mathematics of Arrays (MoA). Companion Papers I-IV formally derive kernels for attention forward, backward, fused forward+backward, decode, and the complete block (RMSNorm, gated MLP) as a hardware-independent specification (DNF) transformed to a machine-specific realization (ONF) via gamma, with verification to machine precision against PyTorch. This paper checks those predictions against measured performance on two HPC clusters (Purdue Anvil, NCSA Delta) across CPU and GPU. Three results stand out. (1) We identify and fix a GPU regression: fusing forward+backward, proven to avoid materializing an O(n^2) intermediate, initially ran slower than naive on GPU due to atomic contention. Profiling confirmed 2.00x more atomic instructions; a targeted ONF rewrite reversed it, yielding up to 2.5x speedup. (2) Identical derivations produce markedly different real costs by topology: 535x NUMA-locality penalty on one cluster vs <3x oversubscription on another, showing optimal deployment is a function of the machine's array structure. (3) We report a partially resolved anomaly: identical denotational computations run faster in C than Fortran on CPU but faster in Fortran than C on GPU, narrowed to one dominant kernel and one memory-latency stall mechanism (3.17x time gap matches 3.35x stall gap). We treat hardware-specific optimization as a routine ONF rewrite with fixed, verified DNF, a candidate methodology for scaling AI onto evolving hardware without re-deriving correctness.