Programmable world model benchmark tests video model rule following
ROWBench: Do Video Models Render What the Program Specifies?
Computer Vision and Pattern Recognition
Summary
It is hard to tell if video-based AI models truly understand and follow all the details of a programmed world. The authors created PROWBench, a set of many short videos with detailed records of what should happen in the scene. This lets them check if a video model’s outputs really match the program’s rules and events, even ones outside the camera view. They also built tools to generate and view these scenes from different perspectives and measure how well models follow instructions and display interactions.
What this means in practice
- •For game engine developers: Evaluate if video models faithfully reflect programmed game world rules and interactions in visual outputs to improve engine reliability.
- •For augmented reality developers: Test AR video models’ alignment to scripted dynamic events to ensure consistent rendering of virtual elements with planned scenarios.
Authors
Zheng-Hui Huang, Guixu Lin, Yu-Ju Tsai, Jian-Kai Zhu, Fengbo Lan, Yu-Lun Liu, Yung-Yu Chuang, Kaipeng Zhang, Zhixiang Wang
Abstract
Programmable world models separate executable dynamics from visual generation, offering a promising foundation for next-generation game engines. However, their visual adherence to explicit rules and interactions remains insufficiently evaluated. Existing benchmarks assess visual quality, controllability, and instruction or physical adherence, but rarely test fidelity to fine-grained, program-specified world events. We introduce PROWBench, comprising 170 programmatically constructed episodes and 600 proxy videos covering diverse scenes and interactions. PROWBench logs entity states and timestamped events, including those outside the camera's field of view, as replayable world records, from which it renders synchronized views and proxy representations. This enables generated videos to be checked against the observable consequences of program execution. An extensible framework constructs scenes, controls behaviors, and can render each camera view in different representations, such as coarse 3D, and bounding boxes. The benchmark covers first- and third-person perspectives, with synchronized multi-view observations available for a subset of episodes. Grounded in these records, PROWBench evaluates entity control, long-horizon memory, and, with two VLM-based metrics, Logic-Render Alignment and Interaction Success Rate, adherence to the prescribed timeline and the visual realization of timestamped engine-recorded events.