TileBench compares tile-based AI programming models for GPU speed
TileBench: A Controlled Benchmark for Performance Evaluation and Bottleneck Diagnosis of Tile-Based Programming Models
Summary
Building fast programs for GPUs is tricky because different programming models behave differently depending on the task. The authors created TileBench, a test set of 45 AI-related tasks with matching implementations in two popular tile-based programming models, Triton and cuTile. They measured how each model performed on NVIDIA B200 GPUs and found that cuTile works better for some specialized tasks, while Triton performs better on many others, especially those with irregular or bandwidth-heavy operations. They also tested kernels generated by language models and saw Triton was more efficient in refining code. TileBench offers a fair way to compare these models for anyone developing GPU programs.
What this means in practice
- •For gpu kernel developers: Evaluate and optimize kernel code by comparing Triton and cuTile performance under realistic AI workloads to choose the best tool.
- •For high-performance computing teams: Diagnose performance bottlenecks in AI kernels running on NVIDIA GPUs using TileBench’s profiling and diagnostics tools.