Papers for
compiler and runtime developers
Papers whose findings have a practical use for this group, as judged from the abstract. Open a paper to read what it means in practice.
Memory optimal transformer kernels show varied speed on different HPC hardware
Validating Memory-Optimal Transformer Kernels on Real Hardware: From Formal Derivation to Measured Performance Across Two HPC Clusters
Abstract: We validate memory-optimal cost functions for transformer kernels derived via the Mathematics of Arrays (MoA). Companion Papers I-IV formally derive kernels for attention forward, backward, fused forward+backward, decode, and the complete block (RMSNorm, gated MLP) as a hardware-independent specification (DNF) transformed to a machine-specific realization (ONF) via gamma, with verification to machine precision against PyTorch. This paper checks those predictions against measured performance on two HPC clusters (Purdue Anvil, NCSA Delta) across CPU and GPU. Three results stand out. (1) We identify and fix a GPU regression: fusing forward+backward, proven to avoid materializing an O(n^2) intermediate, initially ran slower than naive on GPU due to atomic contention. Profiling confirmed 2.00x more atomic instructions; a targeted ONF rewrite reversed it, yielding up to 2.5x speedup. (2) Identical derivations produce markedly different real costs by topology: 535x NUMA-locality penalty on one cluster vs <3x oversubscription on another, showing optimal deployment is a function of the machine's array structure. (3) We report a partially resolved anomaly: identical denotational computations run faster in C than Fortran on CPU but faster in Fortran than C on GPU, narrowed to one dominant kernel and one memory-latency stall mechanism (3.17x time gap matches 3.35x stall gap). We treat hardware-specific optimization as a routine ONF rewrite with fixed, verified DNF, a candidate methodology for scaling AI onto evolving hardware without re-deriving correctness.
Tempo improves processor instruction ordering with lightweight tags
TEMPO: A Tag-Based Framework for Efficient Memory Ordering
Abstract: Weak-memory processors rely on ordering instruc- tions for correctness, yet conventional implementations often en- force them more conservatively than the memory model requires. This over-enforcement manifests as drain-induced retirement stalls at ordering instructions and conservative squash/replay of speculative loads, suppressing legal executions and reducing throughput. We present TEMPO, a tag-based framework for precise microarchitectural implementation of ordering instructions. TEMPO assigns lightweight ordering tags to instructions and decomposes enforcement across retirement-time predicates and completion-time store ordering, allowing the core to enforce required ordering without conservative retirement serialization. TEMPO eliminates unnecessary retirement serialization at ordering instructions and speculative-load squash/replay. In our evaluation, TEMPO reduces geometric-mean normalized exe- cution cycles by 7.9% on native four-thread workloads and improves geometric-mean IPC by 15.9% on an instrumented SPEC2017 dynamic binary translation (DBT) proxy for cross- ISA execution (e.g., x86-on-Arm), while adding only 262 bytes per core.