FlowTool speeds up image retouching by directly predicting editing settings
FlowTool: Controlling Tool Parameter in Image Retouching via Flow Matching
Computer Vision and Pattern Recognition
Summary
Changing pictures to look better is usually done step-by-step by computers using big language models. This paper shows a new way, called FlowTool, that predicts all editing choices at once instead of in order. FlowTool learns from examples how to create good editing plans quickly by transforming random noise into detailed instructions. The authors found that it works better and much faster than previous methods while using less computer memory.
What this means in practice
- •For photo editing software developers: Enhance image editing tools by integrating a faster, more memory-efficient parameter prediction method that improves editing quality and reduces user wait time.$Commercial implications: This paper enables faster and better automated editing parameter generation, allowing photo software makers to build more responsive, higher-quality retouching features.
- •For graphic design teams: Use a tool that quickly generates precise editing settings from instructions, streamlining workflows when adjusting images for marketing or media production.
Authors
Thanh-Long V. Le, Steven Walton, Seunghyun Yoon, Branislav Kveton, Trung Bui, Eunho Yang, Viet Lai
Abstract
Tool-based image editing (image retouching) is commonly formulated with autoregressive multimodal large language models (MLLMs) that sequentially generate reasoning, tool selections, and parameter values. In this work, we present a novel approach to tool-based image editing by framing the task as a flow matching problem. We introduce FlowTool, a framework that directly models the distribution of high-quality tool parameters conditioned on the input image and user instruction using conditional rectified flow. FlowTool combines a vision-language model backbone for multimodal understanding with a Diffusion Transformer parameter generator that transforms Gaussian noise into an editing plan. We train FlowTool with a two-stage supervised flow-matching curriculum, followed by reward-based post-training. Across MMArt-Bench, FlowTool-Eval, ArtEdit-Bench, and MIT-Adobe5K, FlowTool achieves significantly stronger reference-based performance than specialized MLLM editing agents and proprietary MLLMs, while remaining competitive with proprietary models under reference-free evaluation. Moreover, FlowTool significantly improves inference efficiency, reducing latency by at least $50\times$ while requiring nearly $2\times$ less memory than the compared baselines. These results demonstrate that tool-based image editing can be effectively modeled as conditional generation over structured continuous editing parameters, without autoregressive reasoning.