Verifiable visual rewards improve image generation accuracy and generalization

Verifiable Visual Rewards Transfer from Synthetic Scenes to Natural Prompts

Artificial IntelligenceComputer Vision and Pattern Recognition

Summary

Following detailed instructions in generating images, such as getting the number and placement of objects right, is hard because current ways to judge success are unreliable. The authors created a new approach called Verifiable Visual Rewards (VVR), which uses simple geometric scenes where they know exactly what the correct image looks like. They made a large set of tasks and showed that training models with these clear rewards helps models follow complex instructions better, even on real-world prompts. This method improves the quality and consistency of generated images as measured by both automated tests and human preferences.

What this means in practice

Authors

Shuyue Stella Li, Xiaochuang Han, Yulia Tsvetkov, Luke Zettlemoyer

Abstract

Precise instruction following in image generation, such as satisfying object counts and spatial relations, remains an open challenge at least in part because it is learned using unreliable reward models such as object detectors and vision-language models. We introduce Verifiable Visual Rewards (VVR), the first framework for programmatically verifiable image rewards, and show that training on it generalizes to natural prompts. Each VVR task is a scene of geometric objects and relations among them, from which we derive both the prompt and a deterministic verifier, so tasks can be generated in any number and at any chosen complexity. We release VVRBench, with 10,000 tasks over 32 constraint types, and VVRBench-Challenge, with 720 more complex tasks; the strongest model we evaluate---GPT-Image-2.5---solves 21.4% of VVRBench-Challenge. Using VVR scores as rewards for reinforcement learning (RLVVR) raises the accuracy of Stable Diffusion 3.5 Medium on VVRBench from 2.8% to 28.3% and demonstrates consistent easy-to-hard generalization. These gains extend to out-of-domain benchmarks, and mixing VVR into existing objectives further improves overall performance and human preference, motivating the adoption of VVR into standard image generation post-training recipes.