Papers for

gpu library developers

Papers whose findings have a practical use for this group, as judged from the abstract. Open a paper to read what it means in practice.

Hardware aware features improve gpu kernel selection accuracy

Hardware-Aware Features for CUTLASS Kernel Selection

Abstract: GPU libraries such as CUTLASS expose tens of thousands of semantically equivalent kernels for a single operation, making exhaustive autotuning expensive and execution-free selection difficult. Existing analytical selectors require hand-designed performance rules, while learned selectors operate on raw configuration parameters and must infer hardware consequences from data. We introduce a hardware-aware representation for CUTLASS kernel selection that augments candidate configurations with statically computable estimates of induced hardware behavior. We construct a dataset of 4.9 million CUTLASS kernels and train gradient-boosted and neural learning-to-rank models to rank candidates within each problem. On held-out exhaustive evaluation problems, hardware-aware representations reduce selection regret by up to 40\% relative to structural baselines and 64.2\% relative to NVIDIA's matrix-multiply heuristics. We further evaluate data-efficient cross-precision and epilogue-fusion transfer within CUTLASS GEMM, showing that explicitly representing candidate-induced hardware behavior provides a useful inductive bias for learned kernel selection.

Mon 28 SeptMachine LearningPerformance
The gist
Choosing the best way to run certain GPU operations is hard because there are many similar options to pick from. The authors propose adding detailed computer hardware behavior estimates to each choice to help learning algorithms make better selections. They trained models on millions of examples and showed that including these hardware details led to much better choices than older methods. This approach also works well when switching between different computation precisions or operation styles.
Open → 2609.35587v1