KREX boosts GPU kernel benchmarking by sharing access safely
KREX: Concurrent Kernel Benchmarking on Shared GPUs via Region-Granular Exclusivity
Distributed, Parallel, and Cluster ComputingOperating SystemsPerformance
Summary
Measuring how fast GPU programs run often wastes time because the GPU is held exclusively for each test, even when it isn’t needed all the time. The authors created KREX, a system that lets test programs share the GPU except during the most important timing parts. This way, many tests happen together without messing up the timing accuracy. KREX uses clever tricks like pausing other processes and managing CPU cores to keep tests reliable while making much better use of the GPU.
What this means in practice
- •For gpu performance engineers: Run multiple kernel performance tests concurrently with precise timing to speed up GPU optimization workflows.
- •For cloud infrastructure operators: Increase GPU utilization during benchmarking by safely sharing resources without sacrificing timing accuracy.
Authors
Tianyu Feng, Haoxuan Yu, Tianyuan Wu, Lingyun Yang, Daocheng Ying, Yuxiao Wang, Ruibo Fan, Yinghao Yu, Guodong Yang, Liping Zhang, Wei Wang
Abstract
LLM agents automate GPU kernel optimization by repeatedly composing candidates and measuring their duration on real GPUs. Existing systems preserve measurement fidelity by reserving a GPU for an entire agent session or benchmarking command. However, this results in poor utilization because only a small fraction of command execution requires exclusive GPU access. Sharing GPUs could recover this idle capacity, but introduces contention that compromises measurement fidelity and misdirects the agent's search. We present KREX, a runtime for concurrent kernel agent benchmarking with region-granular exclusivity. KREX lets agents mark critical regions involving timing-sensitive operations within a benchmarking command. The runtime then enforces exclusivity within marked regions and allows concurrent execution outside them, achieving high throughput while preserving measurement fidelity. To enforce in-region exclusivity, KREX blocks new competing GPU submissions and drains outstanding work before freezing sibling processes and isolating CPU cores, protecting both GPU execution and the host threads that drive measurements. To maximize off-region concurrency, KREX reuses GPU contexts in persistent context processes to avoid repeated, node-wide serialized context creation. We evaluate KREX on NVIDIA and AMD GPUs. Compared with command-granular exclusivity baselines, KREX delivers up to $3.4\times$ the benchmarking throughput with a negligible p95 timing inflation of $0.30\%$, $1.58\%$, and $3.90\%$ for kernels longer than 10 ms, 1 ms, and 0.1 ms, respectively.