Gpu utilization for large language model inference breaks down efficiency factors

Dissecting GPU Utilization for LLM Inference on Nvidia Hopper

PerformanceHardware ArchitectureDistributed, Parallel, and Cluster ComputingMachine Learning

Summary

Measuring how busy a GPU is during large language model (LLM) use can be misleading because a single percentage doesn't show all the details. The authors explain that during tasks like generating text (decoding), GPU operations handle small pieces inefficiently, leading to underused GPU power. They studied Nvidia's Hopper GPU running models with different settings and used detailed counters to explain why the GPU isn't fully efficient. This helps reveal what exactly limits GPU performance during these tasks.

What this means in practice

  • For ai infrastructure engineers: Optimize GPU resource allocation for LLM inference by using detailed utilization profiles beyond simple percentages to understand performance bottlenecks.
  • For machine learning system builders: Design inference pipelines for LLMs by selecting kernel operations and batch sizes informed by fine-grained GPU utilization insights on Nvidia Hopper GPUs.

Authors

Mohammad Siavashi, Gerald Q. Maguire, Dejan Kostic, Marco Chiesa

Abstract

A single SM utilization percentage can make an LLM inference workload look compute-saturated while hiding how much useful work is being done. The problem is not that the counter is wrong, but that it collapses several different mechanisms into one number. This is most severe during decode, where each request contributes only one new token and dense projection GEMMs become small-row matrix multiplications. On Hopper, the bfloat16 GMMA path executes these operations in fixed 64-row matrix fragments, so small-batch decode can fill only a small fraction of each fragment with real token rows. In this paper, we profile vLLM with FlashAttention-3 and cuBLASLt on an H100 NVL across cold prefill, warm prefill, and decode, sweeping sequence length and batch size. We replace the usual single utilization number with eight counter-validated views derived from raw Nsight Compute reports, each pinned to an NCU counter or explicit formula. Together, these views map utilization gaps to concrete mechanisms - fragment fill, occupancy limits, stall signatures, wave quantization, and kernel selection - across four production models and six per-layer kernel roles.