FP64 matrix multiplication speed improved using FP8 tensor cores on NVIDIA GPUs

Ozaki 2.5: Engineering the Deconstruction Path of fp64-Emulated Dense Matrix Multiplication on FP8 Tensor Cores

Mathematical SoftwareHardware ArchitectureDistributed, Parallel, and Cluster ComputingPerformance

Summary

High-precision calculations called FP64 matrix multiplications are usually slow on modern GPUs. The paper shows how using lower-precision FP8 tensor cores with a special mathematical trick called residue decomposition can speed this up. The authors model how to best break down data and design hardware features that let these conversions happen without slowing down the main calculations. Their approach nearly doubles the effective speed on certain NVIDIA GPUs by removing bottlenecks in data conversion. This work helps improve performance in scientific computing tasks that rely on these calculations.

FP64 matrix multiplicationFP8 tensor coresresidue number systeminteger pipelinesNVIDIA Rubin GPUtensor-memory equilibriummodulus co-designasynchronous copyhigh-performance computingHPL benchmark

Authors

Satoshi Matsuoka

Abstract

FP8 Ozaki II emulates FP64 matrix multiplication by tensor-core products over a CRT residue system; converting the operands into residue planes (the deconstruction term in the Tensor-Memory Equilibrium model of the companion paper "FP8 is All You Need, Part 1") costs integer-pipe and memory resources before tensor instructions issue. This paper engineers that path; every result is a model projection pending measurement. First, a deconstruction-aware model: on the NVIDIA Rubin GPU the emulated rate reaches the arithmetic roof $P_{\rm FP8}/(3r+1)$ ($\approx 473$ TFLOPS at $r=12$) only within one thread-block cluster; larger outputs are re-split on the fly and held at a floor of $\approx 235$ TFLOPS (half the roof, a ratio of three design integers, not a fit), while real solvers' tall/skinny shapes stay near the crossover, $1.6$-$1.9\times$ over simple deconstruction today. Second, the method: convert-once residue workspaces, an exact two-limb constant-reduction GEMM on integer tensor pipes (or pure-SIMT dp4a), and conversion pipelined behind the MMAs, moving the crossover from $\approx 1211$ to $\approx 480$-$730$. Third, modulus co-design: all-byte and hybrid sets, two supply bounds and a carry-corrected E4M3 split of tail moduli. Fourth and central, the closed-form floor names its hardware escape, and the prize is Rubin's: a stream-side residue-conversion mode on the asynchronous copy path (Option C), a narrow fixed-function block sized as a bill of materials, takes plane formation off the arithmetic pipes and lifts the floor from 235 TFLOPS to the full 473-TFLOPS roof at unchanged cluster reach, about doubling HPL-class FP64 per Rubin GPU, and unbinds conversion-bound sparse kernels. The NVIDIA GB300 GPU, whose 135-TFLOPS roof sits at its own floor, gains little; floor and remedy are Rubin-scale. Application traces ground the analysis; constants are script-checked.