Papers for

multi-gpu system engineers

Papers whose findings have a practical use for this group, as judged from the abstract. Open a paper to read what it means in practice.

Argus automates gpu performance tracking across code regions

Argus: Orchestrating Cross-Layer GPU Performance Measurements around Semantic Regions

Abstract: GPU developers and automated optimizers need performance evidence for semantic code regions--such as neural-network operator implementations and pipeline stages--but this evidence is fragmented across profiling tools. Answering a region-level question can require manually constructing probes and program variants, isolating interfering measurements, and mapping evidence to regions and execution contexts. We present Argus, a region-centric measurement planner and runtime that automates this workflow. Clients identify regions with boundary markers and select signals and execution scopes. Argus preserves region identity across compilation, execution, and measurement variants, constructs interference-aware multi-run plans, and orchestrates transformations and profiling across backends. It joins compiler-, hardware-, and system-level evidence using region identity and dynamic execution context, producing reports that record measurement origins and attribution ambiguity. We evaluate Argus across agentic kernel optimization, persistent megakernel optimization, and cross-level PGO. Across 44 persistent-GEMM and attention configurations, Argus improves 39/44 cases and raises AlphaEvolve's geometric-mean speedup from 5.4% to 8.9%. On a persistent TinyLlama-1.1B decode megakernel, an optimization agent reaches 1.65 ms/token with Argus versus 4.92 ms/token without it, producing a kernel $2.1\times$ faster than PyTorch with CUDA Graphs. Finally, Argus-guided cross-level PGO improves compute--communication overlap, increasing throughput by 7% on average across five multi-GPU settings.

Fri 11 SeptDistributed, Parallel, and Cluster ComputingPerformance
The gist
Measuring how different parts of a program run on GPUs is tricky because the information is scattered and hard to get. The authors created Argus, a tool that automatically tracks and organizes performance details for specific code sections on GPUs. Argus helps developers understand performance better by connecting data from different tools and showing how the GPU runs each part. This leads to faster and more efficient programs, especially for complex tasks like neural networks.
Open → 2609.12299v1