Verifiable visual rewards improve image generation accuracy and generalization
Verifiable Visual Rewards Transfer from Synthetic Scenes to Natural Prompts
Artificial IntelligenceComputer Vision and Pattern Recognition
Summary
Following detailed instructions in generating images, such as getting the number and placement of objects right, is hard because current ways to judge success are unreliable. The authors created a new approach called Verifiable Visual Rewards (VVR), which uses simple geometric scenes where they know exactly what the correct image looks like. They made a large set of tasks and showed that training models with these clear rewards helps models follow complex instructions better, even on real-world prompts. This method improves the quality and consistency of generated images as measured by both automated tests and human preferences.
What this means in practice
- •For image generation developers: Improve accuracy of generated images on complex instructions by incorporating verifiable visual reward training.
- •For digital content creators: Produce images that better follow detailed creative instructions by using models trained with verifiable visual rewards.
Authors
Shuyue Stella Li, Xiaochuang Han, Yulia Tsvetkov, Luke Zettlemoyer
Abstract
Precise instruction following in image generation, such as satisfying object counts and spatial relations, remains an open challenge at least in part because it is learned using unreliable reward models such as object detectors and vision-language models. We introduce Verifiable Visual Rewards (VVR), the first framework for programmatically verifiable image rewards, and show that training on it generalizes to natural prompts. Each VVR task is a scene of geometric objects and relations among them, from which we derive both the prompt and a deterministic verifier, so tasks can be generated in any number and at any chosen complexity. We release VVRBench, with 10,000 tasks over 32 constraint types, and VVRBench-Challenge, with 720 more complex tasks; the strongest model we evaluate---GPT-Image-2.5---solves 21.4% of VVRBench-Challenge. Using VVR scores as rewards for reinforcement learning (RLVVR) raises the accuracy of Stable Diffusion 3.5 Medium on VVRBench from 2.8% to 28.3% and demonstrates consistent easy-to-hard generalization. These gains extend to out-of-domain benchmarks, and mixing VVR into existing objectives further improves overall performance and human preference, motivating the adoption of VVR into standard image generation post-training recipes.