Xtrace improves GPU kernel tracing with less slowdown and better accuracy
Xtrace: High-Fidelity GPU Intra-Kernel Tracing via Binary-Level Instruction Splicing
Hardware Architecture
Summary
Tracing GPU programs helps developers understand and improve performance, but existing tools slow down the program too much and sometimes trace different code than what actually runs. The authors created Xtrace, which inserts tracing instructions directly into the compiled program without interfering with its optimization. This approach preserves nearly all of the original code, adds minimal slowdown, and works on many GPU models. Their tests showed Xtrace traces faster and more accurate than previous tools, helping to speed up AI-related GPU programs.
What this means in practice
- •For gpu performance engineers: Profile complex GPU kernels with minimal slowdown and accurate runtime state information across major GPU architectures.
- •For gpu optimization tool developers: Integrate binary-level tracing into GPU profiling tools to improve accuracy and reduce interference during kernel execution.
Authors
Zhuobin Huang, Kai Zhang, Weihao Cui, Hongshi Tan, Liang Luo, Christopher Dewan, Shen Li, Bingsheng He
Abstract
Modern GPU kernels fuse increasingly more work into a single kernel, and intra-kernel tracing has become the mainstream method to profile them. Tracing inserts probes into the kernel to record its runtime states, and the fidelity of the trace determines the efficiency of performance optimization. Unfortunately, existing tools insert probes before compilation. These tools interfere with the compiler's optimizations, so they trace a different binary from the one the GPU executes. They also add significant runtime overhead. Xtrace is the first GPU kernel tracing system with near-zero compile-time interference and minimized runtime overhead. Xtrace inserts probes directly into the compiled kernel binary. It reuses only the registers that hold dead values at the insertion address and resolves all hazards with the compiler's hazard tables. It further schedules the instruction order, register allocation, and control bits to minimize the runtime overhead the probe introduces. Xtrace supports 19 NVIDIA and AMD GPU architectures, and is publicly available for use at https://g-watch.github.io. We evaluate Xtrace on major production large language model (LLM) kernels against the state-of-the-art tracers Neutrino and IKET from NVIDIA. On H100, B300, and MI300X GPUs, Xtrace preserves 94-98% of the instructions of the kernel, while existing tools preserve only 8-48%. Xtrace adds only 0.9-2.8% overhead, while existing tools add 3.8-75.6%. Xtrace guides a coding agent to reach the same FlashAttention-3 performance with 3.9x fewer iterations than existing traces do. Thanks to our binary-level instrumentation, Xtrace also traces the faster closed-source cuDNN kernel, which guides the agent to lift the open-source FlashAttention-4 by 5.2-13.3% in throughput.