Gradient-based method improves image focus in reinforcement learning alignment
SGA-Flow-GRPO: Spatial Gradient-Guided Credit Assignment for Flow-GRPO
Computer Vision and Pattern Recognition
Summary
Aligning AI-generated images with human preferences can be tricky because existing methods treat all parts of the image equally. The authors propose a way to guide AI to focus on important image patches by using detailed gradient information. Their method better assigns credit to different parts of the image during learning, leading to improved alignment with user prompts. Tests show their approach works better than previous methods without needing extra critic models.
What this means in practice
- •For ai model developers: Train generative image models that better match human preferences by using fine-grained spatial credit assignment in reinforcement learning.
- •For digital artists: Use AI tools improved by spatial gradient methods to more accurately generate images aligned with creative prompts.
Authors
Yunkai Yang, Yudong Zhang, Xinying Chen, Bin Luo, Jienan Lyu, Kunquan Zhang, Weitao Wan, Runmin Dong
Abstract
Reinforcement Learning (RL) has proven effective in aligning flow-based generative models with human preferences. Recently, Flow-GRPO has emerged as an efficient critic-free paradigm by calculating advantages over sampled candidate trajectories. However, standard Flow-GRPO applies a uniform scalar advantage across both temporal denoising steps and spatial latent dimensions, without explicitly accounting for the spatial structure of generated images, which may lead to sub-optimal policy updates. To address this, we propose a novel gradient-guided spatial credit assignment framework tailored for Diffusion Transformers (DiTs). We first reformulate the transition-level log-likelihood in Flow-GRPO into a token-wise representation natively aligned with DiT patch architectures, constructing spatially fine-grained importance sampling ratios. To allocate localized credit without rigid, boundary-sensitive segmentation heuristics, we introduce a continuous spatial credit map derived from reward gradients. Crucially, we employ an outlier-robust normalization scheme based on Median Absolute Deviation (MAD) coupled with temperature scaling, effectively eliminating gradient noise while highlighting functional prompt-aligned regions. Extensive evaluations on GenEval show that our approach delivers SOTA alignment quality, achieving a convergence rate comparable to top-tier methods like DiffusionNFT while substantially improving upon Flow-GRPO-based methods in alignment performance.