Papers for

image generation developers

Papers whose findings have a practical use for this group, as judged from the abstract. Open a paper to read what it means in practice.

Verifiable visual rewards improve image generation accuracy and generalization

Verifiable Visual Rewards Transfer from Synthetic Scenes to Natural Prompts

Abstract: Precise instruction following in image generation, such as satisfying object counts and spatial relations, remains an open challenge at least in part because it is learned using unreliable reward models such as object detectors and vision-language models. We introduce Verifiable Visual Rewards (VVR), the first framework for programmatically verifiable image rewards, and show that training on it generalizes to natural prompts. Each VVR task is a scene of geometric objects and relations among them, from which we derive both the prompt and a deterministic verifier, so tasks can be generated in any number and at any chosen complexity. We release VVRBench, with 10,000 tasks over 32 constraint types, and VVRBench-Challenge, with 720 more complex tasks; the strongest model we evaluate---GPT-Image-2.5---solves 21.4% of VVRBench-Challenge. Using VVR scores as rewards for reinforcement learning (RLVVR) raises the accuracy of Stable Diffusion 3.5 Medium on VVRBench from 2.8% to 28.3% and demonstrates consistent easy-to-hard generalization. These gains extend to out-of-domain benchmarks, and mixing VVR into existing objectives further improves overall performance and human preference, motivating the adoption of VVR into standard image generation post-training recipes.

Mon 28 SeptArtificial IntelligenceComputer Vision and Pattern Recognition
The gist
Following detailed instructions in generating images, such as getting the number and placement of objects right, is hard because current ways to judge success are unreliable. The authors created a new approach called Verifiable Visual Rewards (VVR), which uses simple geometric scenes where they know exactly what the correct image looks like. They made a large set of tasks and showed that training models with these clear rewards helps models follow complex instructions better, even on real-world prompts. This method improves the quality and consistency of generated images as measured by both automated tests and human preferences.
Open → 2609.35641v1

Unbalanced optimal transport improves one-step image generation quality

One-Step Generative Modeling via Unbalanced Optimal Transport

Abstract: Drifting models enable one-step generation by amortizing distribution transport into training, but this efficiency places greater demands on the transport field estimated at each update. In large-scale training, the field is computed from finite mini-batches of generated and real samples, which provide only imperfect approximations to the underlying distributions. Balanced optimal transport enforces exact mass matching within every mini-batch, making the estimated field sensitive to the particular composition of the real-data batch. We find that generated and real samples should be treated asymmetrically: letting the mass assigned to real samples adapt while keeping every generated sample fully transported improves generation across six feature-space metrics in controlled ablations, and is more robust to the relaxation strength than relaxing both marginals simultaneously, which falls below balanced transport under stronger relaxation. Motivated by this observation, we propose Unbalanced Optimal Transport Gradient Flow (UOT-GF), which keeps the generated-sample marginal fixed and relaxes only the real-data marginal. Under identical settings at DiT-B/2 on ImageNet-256, UOT-GF improves Fréchet Inception Distance (FID) from 1.53 to 1.46 over the balanced W-Flow baseline; scaling the same recipe yields 1.34 and 1.22 FID at L/2 and XL/2, the best FID among the one-step models we compare. We further derive the induced UOT transport force, establish a kinetic Vlasov--Fokker--Planck formulation whose overdamped zero-temperature limit recovers the drifting dynamics, and characterize non-target stationary states together with sufficient conditions for convergence.

Sat 26 SeptInformation TheoryMachine Learning
The gist
Generating images in a single step is faster but can be tricky because the model tries to match two sets of images exactly, which can cause mistakes during training. The authors found that allowing the real images’ importance to change while keeping generated images fully accounted for leads to better picture creation. They designed a method called Unbalanced Optimal Transport Gradient Flow (UOT-GF) that uses this idea and showed it improves image quality on ImageNet. They also studied the math behind this method to understand why it works and when it succeeds.
Open → 2609.32708v1

Adaptive joint attention speeds up conditional image generation with diffusion transformers

RefAdapt-DiT: Adaptive Joint Attention for Reference-Conditioned Diffusion Transformers

Abstract: Diffusion Transformers (DiTs) have become the standard backbone for high-quality generative modeling, yet deploying them in conditional generation tasks remains computationally prohibitive because bidirectional joint attention repeatedly processes large reference streams. While existing optimization schemes mitigate generic temporal redundancy, they typically rely on coarse-grained static reuse and overlook the distinct dynamics of references and targets. Specifically, we observe that reference representations often evolve slowly along the generation trajectory, while the target often assigns little attention mass to them; reference drift and this target-to-reference exposure jointly shape how strongly stale reference states affect the target. To exploit these patterns, we introduce \RefAdapt, a training-free framework for adaptive control of joint attention between references and targets. Instead of rigid static strategies, \RefAdapt combines consecutive target-Q change with previously observed target-to-reference attention mass to control reference computation adaptively at block granularity. Under ultra-few-step settings, \RefAdapt enables speedups of up to $2.097\times$ on 4-step MiniMax H3 and $3.54\times$ on 8-step Qwen Image Edit, while maintaining comparable visual quality.

Sat 26 SeptComputer Vision and Pattern Recognition
The gist
Generating images with special AI models called Diffusion Transformers is usually slow when these models need extra information to create images, because they process reference details many times. The authors found that parts of this reference information often change slowly and are sometimes ignored by the model, meaning some computations are unnecessary. They designed a new method called RefAdapt that smartly decides when to update this reference information based on how much the model is paying attention to it, without retraining the AI. This approach speeds up the image generation process significantly while keeping similar image quality.
Open → 2609.32415v1

Compute allocation improves training quality in drifting image models

Refresh or Realize? Compute Allocation in Drifting Models

Abstract: Drifting Models train a one-step generator by recomputing a finite-sample drift field at every iteration and taking an optimizer step toward the drifted target. The field says how generated samples should move, but the step is taken in parameters shared by all samples, so the motion the network actually makes need not match the motion it was given. This leaves a basic training question open: should extra compute go into fitting the current target more closely, or into recomputing the field? We study it on ImageNet 256x256. Holding the target fixed for k optimizer steps and measuring the realized displacement, we find that deeper fitting does bring the network closer to the frozen target, and that the number of steps needed before it makes any net progress drops from about sixteen early in training to one later on. When the extra steps come for free, k=2 also lowers FID. Once they are paid for, the result flips: at approximately matched measured wall-clock, spending the budget on fresh fields gives lower FID than deeper fitting, on both training seeds. The target itself shows why a fresh field is worth so much. Redrawing the finite support rotates its direction far more than a parameter update does (cosine ~0.3-0.6 against ~0.95), and a correction that is optimal in field space is not reliably better in FID than a parameter-free one. For Drifting, fitting each target well and spending compute well are different goals.

Sat 26 SeptMachine Learning
The gist
This paper looks at how to best use computing power when training drifting models, a type of image generator. The models learn by moving images step-by-step based on calculated directions, but the directions change over time, so the model must decide whether to spend more time refining current directions or recalculating new ones. The authors found that early on, more focus on refining helps, but later, recalculating directions more often leads to better image quality. This insight helps balance computation during training to get better final images.
Open → 2609.32298v1

New multi-agent system scores image quality more reliably on open-ended traits

An Evolutionary Agentic Approach for Open-ended Image Quality Perception

Abstract: Generative models are rapidly expanding image quality assessment (IQA) beyond traditional fidelity factors to emerging dimensions such as physical plausibility and text-rendering correctness. However, existing IQA models rely on fixed definitions and heavy supervision, making them difficult to extend to open-ended perceptual dimensions. We identify holistic bias as an important limitation: when scoring an unseen dimension, models reuse generic quality priors, leading to scoring errors and rank inversion. To address this, we propose PACE (Perceptual Agentic Collaborative Evolution), a training-free multi-agent framework that formulates open-ended IQA as explicit protocol construction. Given a target dimension, PACE uses collaborative agents to construct an evaluation protocol composed of verifiable Visual Question Answering (VQA) probes, grounding evaluation in concrete visual evidence rather than holistic impressions. The resulting protocol is calibrated using only four human-annotated images per dimension, while a dual-track scoring mechanism aligns model perception with human scoring scales. Across traditional IQA, structural fidelity, context-aware aesthetics, and newly defined open-ended dimensions, PACE consistently improves its MLLM backbone, achieving competitive performance across diverse IQA settings, and reduces the Holistic Override Rate (HOR) from 44.4\% to 8.6\%.

Sat 19 SeptComputer Vision and Pattern RecognitionArtificial Intelligence
The gist
Assessing the quality of images involves more than just measuring sharpness or fidelity—it includes things like how natural or correct text in images looks. Existing methods struggle when asked to judge new kinds of quality because they rely on fixed ideas and lots of training. The authors developed PACE, a method where multiple AI agents work together to create clear, question-based tests instead of relying on broad impressions. This leads to better alignment with how people judge image quality on many different traits, even ones not seen before.
Open → 2609.22942v1

Hardware accelerator speeds up diffusion transformer machine learning models

The World Model Hardware Accelerator

Abstract: Diffusion transformers invert the arithmetic that autoregressive decoding made familiar. There is no token-by-token recurrence: every denoising step is a full-sequence forward pass over static shapes, so the entire schedule is known at compile time and the only serial dimension is the step count itself. We exploit that structure in WMHA, a latency-first diffusion-transformer inference accelerator: a very-long-instruction-word sequencer issues four engines from one instruction word, a weight-stationary 16x16 dual-dot array streams FP8 and BF16 contractions, and a single-pass online-softmax attention pipeline keeps keys and values resident through a skewed software pipeline. The design is specified in a frozen micro-architecture document, implemented in synthesizable SystemVerilog, and verified against a double-precision reference model by a UVM environment whose acceptance criterion is semantic: the device must run a real denoising trajectory and reduce mean squared error against a clean latent by at least a factor of ten. It does so by a factor of 23, at both synthesized configurations, with zero element failures across 237 million checked values. Eleven application benchmarks built from published model shapes, including the original diffusion-transformer configuration, run on the device and report measured occupancy beside separately labelled projections. Five engines are taken to routed layout in sky130 with parasitic-annotated timing and measured-activity power; the full chip is synthesized, and the host limit that stopped its place-and-route is quantified together with the machine that would remove it.

Mon 14 SeptHardware Architecture
The gist
Many machine learning models generate data step by step, which can be slow. The authors worked on a type of model called diffusion transformers that process all steps at once instead of one by one. They designed a special computer chip that speeds up these models by making the process more parallel and efficient. The chip was tested carefully and showed great accuracy and performance improvements for various model tasks.
Open → 2609.16244v1