Papers for

gpu kernel developers

Papers whose findings have a practical use for this group, as judged from the abstract. Open a paper to read what it means in practice.

TileBench compares tile-based AI programming models for GPU speed

TileBench: A Controlled Benchmark for Performance Evaluation and Bottleneck Diagnosis of Tile-Based Programming Models

Abstract: Tile-based programming models, such as Triton and cuTile, aim to simplify high-performance kernel development, but their practical performance, tuning behavior, and usability remain difficult to compare systematically. We present TileBench, a controlled benchmark for evaluating Triton and cuTile on NVIDIA B200 GPUs under matched operator semantics and comparable implementation structures. TileBench contains 45 operators covering diverse AI-kernel patterns and memory/computation behaviors. Each task provides a PyTorch reference, verified Triton and cuTile implementations, standardized data-types (dtype) and input-size sweeps, default and autotuned configurations, roofline-based metrics, and profiling-guided diagnosis. Our evaluation shows that performance gaps are workload-dependent: cuTile excels on a small cluster of Tensor-Core/TMA-friendly kernels, while Triton is stronger on many irregular, streaming, and bandwidth-bound operators. We further evaluate LLM-generated cuTile and Triton kernels and find that Triton is consistently more token-efficient than cuTile under the same iterative refinement protocol. TileBench is publicly available at https://github.com/Deep-Learning-Profiling-Tools/Tilebench.

Thu 24 SeptPerformanceProgramming Languages
The gist
Building fast programs for GPUs is tricky because different programming models behave differently depending on the task. The authors created TileBench, a test set of 45 AI-related tasks with matching implementations in two popular tile-based programming models, Triton and cuTile. They measured how each model performed on NVIDIA B200 GPUs and found that cuTile works better for some specialized tasks, while Triton performs better on many others, especially those with irregular or bandwidth-heavy operations. They also tested kernels generated by language models and saw Triton was more efficient in refining code. TileBench offers a fair way to compare these models for anyone developing GPU programs.
Open → 2609.29067v1

Tensor program bugs found faster by checking outputs one location at a time

The Output-Space Hypothesis: Enumerative Equivalence Checking for Tensor Programs

Abstract: Tensor programs, as used in deep learning models, are a prime target for optimization, as small performance improvements can have a large impact across training or inference workloads. However, such optimizations are complicated and can produce subtle bugs. Traditionally, correctness is assumed when differential testing against a reference on random inputs fails to reveal bugs. However, the inputs to these programs are massive tensors, and finding bugs can require generating extremely low likelihood inputs with precise relationships among their values. We propose a novel way to find bugs more consistently by flipping the quantifiers. Rather than generating a single input and checking all output tensor locations for equivalence, what if you could check a single output tensor location's equivalence for all inputs? We implement this idea in a system, \dirigo, by using a novel symbolic execution strategy. We demonstrate that \dirigo can find bugs effectively in a public dataset of 6,988 AI-written CUDA kernels that are all marked correct by differential testing. Of these, \dirigo finds 600 kernels that are actually buggy, and finds 97.3\% of those bugs within two minutes.

Thu 17 SeptProgramming LanguagesMachine LearningSoftware Engineering
The gist
Tensor programs used in AI models are tricky to optimize because small bugs can cause big problems. Traditional tests check many outputs for one input, but this misses rare bugs hidden in complex data. The authors flipped this by checking one output location for all possible inputs, catching subtle mistakes more reliably. Their system, Dirigo, found hundreds of bugs that older tests missed, most within minutes.
Open → 2609.19611v1

Argus automates gpu performance tracking across code regions

Argus: Orchestrating Cross-Layer GPU Performance Measurements around Semantic Regions

Abstract: GPU developers and automated optimizers need performance evidence for semantic code regions--such as neural-network operator implementations and pipeline stages--but this evidence is fragmented across profiling tools. Answering a region-level question can require manually constructing probes and program variants, isolating interfering measurements, and mapping evidence to regions and execution contexts. We present Argus, a region-centric measurement planner and runtime that automates this workflow. Clients identify regions with boundary markers and select signals and execution scopes. Argus preserves region identity across compilation, execution, and measurement variants, constructs interference-aware multi-run plans, and orchestrates transformations and profiling across backends. It joins compiler-, hardware-, and system-level evidence using region identity and dynamic execution context, producing reports that record measurement origins and attribution ambiguity. We evaluate Argus across agentic kernel optimization, persistent megakernel optimization, and cross-level PGO. Across 44 persistent-GEMM and attention configurations, Argus improves 39/44 cases and raises AlphaEvolve's geometric-mean speedup from 5.4% to 8.9%. On a persistent TinyLlama-1.1B decode megakernel, an optimization agent reaches 1.65 ms/token with Argus versus 4.92 ms/token without it, producing a kernel $2.1\times$ faster than PyTorch with CUDA Graphs. Finally, Argus-guided cross-level PGO improves compute--communication overlap, increasing throughput by 7% on average across five multi-GPU settings.

Fri 11 SeptDistributed, Parallel, and Cluster ComputingPerformance
The gist
Measuring how different parts of a program run on GPUs is tricky because the information is scattered and hard to get. The authors created Argus, a tool that automatically tracks and organizes performance details for specific code sections on GPUs. Argus helps developers understand performance better by connecting data from different tools and showing how the GPU runs each part. This leads to faster and more efficient programs, especially for complex tasks like neural networks.
Open → 2609.12299v1