Benchmark tests gpu coding agents on physical simulation tasks

GPUPhysBench: Benchmarking Coding Agents for Correct and Efficient GPU Physics Simulation

Distributed, Parallel, and Cluster ComputingArtificial Intelligence

Summary

Writing fast and correct programs that run on graphics cards to simulate physical systems is very challenging. The authors introduce GPUPhysBench, a set of 50 tasks that checks whether coding AI agents can write code to simulate fluids, solids, and granular materials accurately and efficiently. They test different AI models on these tasks and find that while some can complete all tasks correctly, only a few can run the simulations nearly as fast as expert-written code. The hardest parts involve handling collisions and solving complex equations. This benchmark helps measure both correctness and speed in programming physical simulations on GPUs.

What this means in practice

  • For gpu software developers: Evaluate and improve AI coding agents’ ability to write fast and correct GPU code for complex physics simulations using a standardized benchmark.
  • For video game physics engineers: Test automated code generation tools on real physical simulation tasks to gauge their effectiveness in producing reliable and efficient physics effects.

Authors

Yuchen Sun, Jinjin He, Sinan Wang, Bo Zhu

Abstract

Writing fast GPU code for physical simulation is difficult: implementations must preserve numerical accuracy while handling irregular data access, synchronization, and iterative solvers. We introduce GPUPhysBench, a benchmark of 50 tasks testing whether coding agents can meet these demands. Tasks cover fluids, deformable solids, and granular materials, from individual simulation operators to complete simulators. Agents write, compile, test, and optimize GPU code with access to a NVIDIA GPU under fixed time budgets. We report pass rates and runtime performance relative to expert-optimized reference implementations. In a single-attempt evaluation of six frontier model-harness pairs, the two strongest pass all 50 tasks, but even the fastest reaches at least 0.9 the reference speed on only 22% of them, and no submission is more than 5% faster than the reference. The largest gaps arise in collision detection, constraint solving, and iterative solvers. GPUPhysBench brings physical simulation workloads to coding-agent evaluation, testing both the ability to implement numerical methods correctly and the ability to make them run efficiently.