Papers for

graphic design software developers

Papers whose findings have a practical use for this group, as judged from the abstract. Open a paper to read what it means in practice.

Visual problem solving benchmark improves generative model planning

SolveEdit: Benchmarking Visual Problem Solving in Generative Models

Abstract: Machine intelligence is often evaluated through abstract reasoning problems, yet many real-world problems are visual, such as arranging objects, repairing layouts, or tracing routes. Solving these problems requires understanding a scene, inferring what must change to achieve a goal, and realizing that change without disturbing unrelated content. However, existing benchmarks mainly evaluate perception, generation, or explicitly specified transformations, leaving goal-driven visual problem solving underexplored. To bridge this gap, we introduce SolveEpIT, a benchmark for visual problem solving through scene transformation. Given an image and a goal, a model must infer a valid transformation from the request, the scene, or a visually expressed rule, then execute it while preserving unrelated content. SoLvEEDrr contains 2,728 cases. Atomic transition contracts specify required and protected conditions, enabling SoLvEScoRE to measure completion and unintended changes without a single reference output. The strongest evaluated model achieves only57.0% SolvEScore. We further introduce SolveEdiT-PLAN, a two-stage visual planner that instantiates the transition before generation. Under matched single-generation evaluation, it improves SoLvEScoRE by 9.1 points on average across three tested generators, including a gain from 57.0% to 71.6% for GPT-Image-2, without modifying the editor.

Mon 28 SeptComputer Vision and Pattern RecognitionArtificial Intelligence
The gist
Many real-world problems involve changing images to reach a goal, like moving objects without messing up the rest of the picture. The authors created a test called SolveEpIT to check if computer programs can do these tasks well. They also made a new two-step method that helps programs plan changes before making them, which works better than previous ways. Even the best current program only solved about half the problems correctly, showing this is a hard challenge.
Open → 2609.35504v1

Video text editing improved by aligning glyphs with text movement

Enhanced Video Text Editing with Trajectory-Aligned Glyph Rendering

Abstract: Video text editing aims to replace or add text in a video while keeping the rest of the video unchanged, which requires the edited text to be correct in every frame and to move coherently with the scene. Despite the remarkable progress of video diffusion models, they struggle to reproduce exact stroke structures and often produce garbled or wrong characters, especially for characters with complex strokes. To address this, we propose a trajectory-aligned glyph rendering reference that provides explicit per-frame glyph guidance following the position and perspective of the text, and a depth-normalized recognizer feature supervision that supervises the generated text on multi-depth features of a frozen text recognizer with per-depth normalized errors, targeting stroke errors overlooked by the diffusion loss. We further build VTEdit, a benchmark of 288 real-scene clips with 440 annotated text trajectories covering text replacement and text addition, which will be publicly released to facilitate future research. Experiments on VTEdit show that our method outperforms image text editing methods, video editing methods, and commercial models in text accuracy and background preservation, achieving a sentence accuracy of 0.9408, and receives the highest preference in a user study.

Mon 28 SeptComputer Vision and Pattern Recognition
The gist
Editing text in videos is hard because the new text needs to match the movement and angle of the original in every frame, which existing video diffusion methods struggle with. The authors created a way to guide the editing using exact shapes of letters matched to the text’s path and a new method to check mistakes in letter strokes more carefully. They also made a new set of 288 video clips with annotated text to test such methods. Their approach makes text changes more accurate and keeps the background better than previous methods and commercial tools.
Open → 2609.34178v1

Ruler improves svg image generation with detailed rubric rewards

RULER: Instance-aware Rubric Rewards for SVG Generation

Abstract: Generating Scalable Vector Graphics (SVG) code from natural-language instructions is an open-ended task without absolute visual ground truth, leaving both evaluation and policy optimization without a faithful signal. Scalar metrics (CLIP, Aesthetic) calibrated on natural images transfer poorly to stylized vector content, and reusing them as RL rewards triggers reward hacking. We address both limitations with rubric-based scoring. We first establish empirically that prompting a vision-language judge with a multi-axis rubric correlates with human judgments far better than scalar metrics, both across samples and within instructions. Building on this finding, we introduce RULER (Instance-aware Rubric Rewards for Reinforcement Learning), which converts each instruction into an instance-aware rubric of six items spanning semantic, visual, and stylistic axes; a judge VLM scores rendered rollouts item-by-item, and the weighted satisfactions form a fine-grained reward optimized via Group Relative Policy Optimization. Because the rubric is derived from text alone, RULER requires neither paired SVG ground truth nor human preference labels. On MMSVG-Illustration and MMSVG-Icon, RULER lifts the rubric score from 0.432/0.395 to 0.693/0.683, surpassing dedicated SVG specialists and matching the substantially larger DeepSeek-V3, with ablations identifying rubric design as the active lever for RL on open-ended SVG generation. The project page is available at https://hangyuran.github.io/RULER/.

Mon 21 SeptComputer Vision and Pattern Recognition
The gist
Generating SVG images from text instructions is tricky because there often isn’t one right answer, making it hard to evaluate or improve models. The authors show that using a detailed rubric with multiple criteria works better than simple scores for judging image quality. They created RULER, a system that makes a custom rubric for each instruction and scores generated images on various aspects using a vision-language model, then improves the generation through reinforcement learning based on these scores. This approach improves quality without needing exact example images or human feedback.
Open → 2609.25270v1

Model decodes graphic design elements in order to improve detection

Detect Anything in Graphic Design: Element-Level Rewards for Autoregressive Detection

Abstract: Graphic designs, such as posters, advertisements, and infographics, are an important medium for communicating information and shaping understanding. Unlike natural images, they consist of layered elements with explicit compositional order. However, existing object detection models treat these elements as an unordered set, leaving compositional order unexploited. To address this limitation, we present Detect Anything in Graphic Design (DAD), a model that formulates graphic design detection as compositional deconstruction. It decodes elements in compositional order, using lower-layer elements to better detect higher-layer ones. The key feature of DAD is amodal detection, which predicts the full bounding box of each element, including regions occluded by elements placed above it. Building on this formulation, we propose Element Relative Policy Optimization (EleRPO), which extends GRPO from sequence-level supervision to element-level optimization. EleRPO provides fine-grained training signals that capture how each detected element contributes to overall detection quality, and works synergistically with compositional order to improve detection performance. To support training and evaluation, we build a dataset of 10 million graphic designs. Experiments show that DAD outperforms all baselines and achieves human-level performance in amodal detection, supporting effective image-to-layer decomposition. EleRPO consistently improves over GRPO across nine detection benchmarks.

Mon 7 SeptComputer Vision and Pattern RecognitionMachine Learning
The gist
Graphic designs like posters are made of many layers stacked on top of each other, but current detection tools treat these layers like a jumbled set. The authors created a model, DAD, that looks at elements in the order they appear, using information from lower layers to detect higher ones better. Their system can also predict the full shape of each element, even if part of it is hidden behind others. They also developed a new training method that gives detailed feedback for each element to help the model learn better. They tested their approach on a huge dataset of graphic designs and found it works as well as humans for detecting covered elements.
Open → 2609.07072v1