Visual problem solving benchmark improves generative model planning

SolveEdit: Benchmarking Visual Problem Solving in Generative Models

Computer Vision and Pattern RecognitionArtificial Intelligence

Summary

Many real-world problems involve changing images to reach a goal, like moving objects without messing up the rest of the picture. The authors created a test called SolveEpIT to check if computer programs can do these tasks well. They also made a new two-step method that helps programs plan changes before making them, which works better than previous ways. Even the best current program only solved about half the problems correctly, showing this is a hard challenge.

What this means in practice

  • For graphic design software developers: Create smarter image editing tools that understand user goals and preserve unrelated content during scene adjustments.$Commercial implications: Enables advanced photo and layout editing features for creative software sold to designers and artists.
  • For robotics system engineers: Improve robots’ visual reasoning to rearrange objects and repair layouts based on goal-directed visual input.

Authors

Wenjie Shu, Yexin Liu, Harold Haodong Chen, Xuerui Qiu, Zehan Wang, Yidi Zhang, Yizhan Chen, Zunwei Wang, Minghao Liu, Qi Chen, Harry Yang, Xiaogang Xu

Abstract

Machine intelligence is often evaluated through abstract reasoning problems, yet many real-world problems are visual, such as arranging objects, repairing layouts, or tracing routes. Solving these problems requires understanding a scene, inferring what must change to achieve a goal, and realizing that change without disturbing unrelated content. However, existing benchmarks mainly evaluate perception, generation, or explicitly specified transformations, leaving goal-driven visual problem solving underexplored. To bridge this gap, we introduce SolveEpIT, a benchmark for visual problem solving through scene transformation. Given an image and a goal, a model must infer a valid transformation from the request, the scene, or a visually expressed rule, then execute it while preserving unrelated content. SoLvEEDrr contains 2,728 cases. Atomic transition contracts specify required and protected conditions, enabling SoLvEScoRE to measure completion and unintended changes without a single reference output. The strongest evaluated model achieves only57.0% SolvEScore. We further introduce SolveEdiT-PLAN, a two-stage visual planner that instantiates the transition before generation. Under matched single-generation evaluation, it improves SoLvEScoRE by 9.1 points on average across three tested generators, including a gain from 57.0% to 71.6% for GPT-Image-2, without modifying the editor.