Argus automates gpu performance tracking across code regions
Argus: Orchestrating Cross-Layer GPU Performance Measurements around Semantic Regions
Distributed, Parallel, and Cluster ComputingPerformance
Summary
Measuring how different parts of a program run on GPUs is tricky because the information is scattered and hard to get. The authors created Argus, a tool that automatically tracks and organizes performance details for specific code sections on GPUs. Argus helps developers understand performance better by connecting data from different tools and showing how the GPU runs each part. This leads to faster and more efficient programs, especially for complex tasks like neural networks.
What this means in practice
- •For gpu kernel developers: Measure and optimize code sections in GPU programs automatically to improve speed and efficiency in neural network implementations and other workloads.
- •For multi-gpu system engineers: Increase throughput on multi-GPU setups by using coordinated profiling data to improve compute and communication overlap within programs.
Authors
Jianzhu Yao, Yue Guan, Srivatsan Ramesh, Yuanwei Fang, Jian Jiao, Boda Li, Yueming Hao, Xinwei Qiang, Pramod Viswanath, Yufei Ding, Bill Yoshimi, Alexey Loginov, Shane Nay, Adnan Aziz
Abstract
GPU developers and automated optimizers need performance evidence for semantic code regions--such as neural-network operator implementations and pipeline stages--but this evidence is fragmented across profiling tools. Answering a region-level question can require manually constructing probes and program variants, isolating interfering measurements, and mapping evidence to regions and execution contexts. We present Argus, a region-centric measurement planner and runtime that automates this workflow. Clients identify regions with boundary markers and select signals and execution scopes. Argus preserves region identity across compilation, execution, and measurement variants, constructs interference-aware multi-run plans, and orchestrates transformations and profiling across backends. It joins compiler-, hardware-, and system-level evidence using region identity and dynamic execution context, producing reports that record measurement origins and attribution ambiguity. We evaluate Argus across agentic kernel optimization, persistent megakernel optimization, and cross-level PGO. Across 44 persistent-GEMM and attention configurations, Argus improves 39/44 cases and raises AlphaEvolve's geometric-mean speedup from 5.4% to 8.9%. On a persistent TinyLlama-1.1B decode megakernel, an optimization agent reaches 1.65 ms/token with Argus versus 4.92 ms/token without it, producing a kernel $2.1\times$ faster than PyTorch with CUDA Graphs. Finally, Argus-guided cross-level PGO improves compute--communication overlap, increasing throughput by 7% on average across five multi-GPU settings.