Unified policy optimization improves image diversity and robot task success

Unified Trajectory Matching Policy Optimization: Diverse T2I Generation and VLA Generalization

Computer Vision and Pattern Recognition

Summary

Generating images from text often leads to repetitive pictures because the training focuses only on the best outcome, ignoring variety. The authors found a way to train AI policies that balance getting high rewards with keeping diverse results, so images look different but still good. Their method also helps robots pick different ways to finish tasks, improving adaptability when conditions change. Experiments show this approach outperforms previous methods on both image generation and robot tasks.

What this means in practice

  • For ai image generation developers: Generate more diverse and high-quality images from text prompts by using policy optimization that balances reward and variety.$Commercial implications: Enables production of creative and varied AI-generated images for content creation platforms that compete on image diversity and quality.
  • For robotics engineers: Improve robot task success rates and adaptability by enabling multiple action strategies through feedback-conditioned policy training.

Authors

Zhiyuan Ma, Jiaming Li, Lingzhen Li, Yu Liu, Xuekai Zhu, Dingkang Liang, Kaiyan Zhang, Jianjun Li, Bowen Zhou, Xiang Bai

Abstract

Reward-maximizing reinforcement learning (RL) is widely used to post-train stochastic diffusion and flow policies for text-to-image (T2I) generation. However, reward-maximizing RL causes policy mode collapse even under reference KL or entropy regularization, reducing the policy to a single high-reward mode. In T2I, this produces similar images and reward hacking. When extended to vision-language-action (VLA) models, the same collapse removes alternative successful strategies and weakens task and scene generalization. To address this limitation, we introduce Unified Trajectory Matching Policy Optimization (Uni-TMPO), a unified RL post-training framework for diffusion and flow policies. First, Uni-TMPO converts standardized rewards into a target distribution within each trajectory group and derives the policy distribution from trajectory log probabilities. Then, forward Kullback-Leibler optimization matches the two distributions instead of maximizing expected reward. A progress-conditioned coarse-to-fine scheduler efficiently constructs T2I trajectories. Within the unified framework, feedback-conditioned sampling uses updated observations to construct VLA trajectories. Extensive experiments show that Uni-TMPO achieves higher T2I rewards and VLA ID success rates than the strongest baselines. More importantly, it achieves the best T2I reward-diversity-efficiency trade-off and VLA generalization to held-out tasks and scenes, while real-robot evaluation demonstrates the value of multiple action strategies when the higher-reward target is blocked.