Papers for

ai image generation developers

Papers whose findings have a practical use for this group, as judged from the abstract. Open a paper to read what it means in practice.

Unified policy optimization improves image diversity and robot task success

Unified Trajectory Matching Policy Optimization: Diverse T2I Generation and VLA Generalization

Abstract: Reward-maximizing reinforcement learning (RL) is widely used to post-train stochastic diffusion and flow policies for text-to-image (T2I) generation. However, reward-maximizing RL causes policy mode collapse even under reference KL or entropy regularization, reducing the policy to a single high-reward mode. In T2I, this produces similar images and reward hacking. When extended to vision-language-action (VLA) models, the same collapse removes alternative successful strategies and weakens task and scene generalization. To address this limitation, we introduce Unified Trajectory Matching Policy Optimization (Uni-TMPO), a unified RL post-training framework for diffusion and flow policies. First, Uni-TMPO converts standardized rewards into a target distribution within each trajectory group and derives the policy distribution from trajectory log probabilities. Then, forward Kullback-Leibler optimization matches the two distributions instead of maximizing expected reward. A progress-conditioned coarse-to-fine scheduler efficiently constructs T2I trajectories. Within the unified framework, feedback-conditioned sampling uses updated observations to construct VLA trajectories. Extensive experiments show that Uni-TMPO achieves higher T2I rewards and VLA ID success rates than the strongest baselines. More importantly, it achieves the best T2I reward-diversity-efficiency trade-off and VLA generalization to held-out tasks and scenes, while real-robot evaluation demonstrates the value of multiple action strategies when the higher-reward target is blocked.

Mon 28 SeptComputer Vision and Pattern Recognition
The gist
Generating images from text often leads to repetitive pictures because the training focuses only on the best outcome, ignoring variety. The authors found a way to train AI policies that balance getting high rewards with keeping diverse results, so images look different but still good. Their method also helps robots pick different ways to finish tasks, improving adaptability when conditions change. Experiments show this approach outperforms previous methods on both image generation and robot tasks.
Open → 2609.34688v1