TileBench compares tile-based AI programming models for GPU speed

TileBench: A Controlled Benchmark for Performance Evaluation and Bottleneck Diagnosis of Tile-Based Programming Models

PerformanceProgramming Languages

Summary

Building fast programs for GPUs is tricky because different programming models behave differently depending on the task. The authors created TileBench, a test set of 45 AI-related tasks with matching implementations in two popular tile-based programming models, Triton and cuTile. They measured how each model performed on NVIDIA B200 GPUs and found that cuTile works better for some specialized tasks, while Triton performs better on many others, especially those with irregular or bandwidth-heavy operations. They also tested kernels generated by language models and saw Triton was more efficient in refining code. TileBench offers a fair way to compare these models for anyone developing GPU programs.

What this means in practice

  • For gpu kernel developers: Evaluate and optimize kernel code by comparing Triton and cuTile performance under realistic AI workloads to choose the best tool.
  • For high-performance computing teams: Diagnose performance bottlenecks in AI kernels running on NVIDIA GPUs using TileBench’s profiling and diagnostics tools.

Authors

Bowen Cui, Zhongchun Zhou, Hao Wu, Tejas Ramesh, Junyu Yin, Jialiang Gu, Keren Zhou

Abstract

Tile-based programming models, such as Triton and cuTile, aim to simplify high-performance kernel development, but their practical performance, tuning behavior, and usability remain difficult to compare systematically. We present TileBench, a controlled benchmark for evaluating Triton and cuTile on NVIDIA B200 GPUs under matched operator semantics and comparable implementation structures. TileBench contains 45 operators covering diverse AI-kernel patterns and memory/computation behaviors. Each task provides a PyTorch reference, verified Triton and cuTile implementations, standardized data-types (dtype) and input-size sweeps, default and autotuned configurations, roofline-based metrics, and profiling-guided diagnosis. Our evaluation shows that performance gaps are workload-dependent: cuTile excels on a small cluster of Tensor-Core/TMA-friendly kernels, while Triton is stronger on many irregular, streaming, and bandwidth-bound operators. We further evaluate LLM-generated cuTile and Triton kernels and find that Triton is consistently more token-efficient than cuTile under the same iterative refinement protocol. TileBench is publicly available at https://github.com/Deep-Learning-Profiling-Tools/Tilebench.