Ruler improves svg image generation with detailed rubric rewards
RULER: Instance-aware Rubric Rewards for SVG Generation
Computer Vision and Pattern Recognition
Summary
Generating SVG images from text instructions is tricky because there often isn’t one right answer, making it hard to evaluate or improve models. The authors show that using a detailed rubric with multiple criteria works better than simple scores for judging image quality. They created RULER, a system that makes a custom rubric for each instruction and scores generated images on various aspects using a vision-language model, then improves the generation through reinforcement learning based on these scores. This approach improves quality without needing exact example images or human feedback.
What this means in practice
- •For graphic design software developers: Improve automatic SVG generation tools by optimizing outputs using rubric-based rewards for better alignment with natural language instructions.
- •For digital content creators: Generate higher quality vector images from descriptive text prompts for illustration and icon design workflows.
Authors
Hangyu Ran, Yuhao Zheng, Yingying Zhang, Kevin Qinghong Lin, Han Peng
Abstract
Generating Scalable Vector Graphics (SVG) code from natural-language instructions is an open-ended task without absolute visual ground truth, leaving both evaluation and policy optimization without a faithful signal. Scalar metrics (CLIP, Aesthetic) calibrated on natural images transfer poorly to stylized vector content, and reusing them as RL rewards triggers reward hacking. We address both limitations with rubric-based scoring. We first establish empirically that prompting a vision-language judge with a multi-axis rubric correlates with human judgments far better than scalar metrics, both across samples and within instructions. Building on this finding, we introduce RULER (Instance-aware Rubric Rewards for Reinforcement Learning), which converts each instruction into an instance-aware rubric of six items spanning semantic, visual, and stylistic axes; a judge VLM scores rendered rollouts item-by-item, and the weighted satisfactions form a fine-grained reward optimized via Group Relative Policy Optimization. Because the rubric is derived from text alone, RULER requires neither paired SVG ground truth nor human preference labels. On MMSVG-Illustration and MMSVG-Icon, RULER lifts the rubric score from 0.432/0.395 to 0.693/0.683, surpassing dedicated SVG specialists and matching the substantially larger DeepSeek-V3, with ablations identifying rubric design as the active lever for RL on open-ended SVG generation. The project page is available at https://hangyuran.github.io/RULER/.