Gradient-guided grouping improves reward-weighted video model training

G$^3$-LoRA: Organizing Reward-Weighted Video Data with Gradient-Guided Grouped LoRA

Computer Vision and Pattern RecognitionMachine Learning

Summary

Training video models using various types of reward signals can be hard because different types of feedback might suggest conflicting improvements. The authors study a way to organize training data by grouping similar types of feedback according to how their training updates interact. They propose a new method called G3-LoRA that clusters these groups based on gradient directions, trains specialized adapters for each, and then merges them to improve the overall model. This method leads to better performance compared to mixing all feedback together, though some specialized skills remain challenging.

What this means in practice

  • For video ai developers: Develop improved video generation models by organizing training data into gradient-compatible groups for specialized adapter training and merging.
  • For multimodal ai engineers: Use gradient-guided clustering to better integrate different reward signals in training models combining text and video inputs.

Authors

Jia Song, Wenhow Li, Lichen Bai, Bada Ye, Zeke Xie

Abstract

Post-training foundation video models on heterogeneous reward-weighted data usually assume that all data categories induce compatible updates. This assumption is fragile when categories correspond to different skills, domains, or evaluation dimensions. We study this problem in text-to-video post-training, where VBench2.0 dimensions define data buckets and an external multimodal reward pipeline assigns sample weights. We propose G$^3$-LoRA (Gradient-Guided Grouped LoRA), a data organization procedure that probes category-level gradients induced by reward-weighted video samples, removes the shared global update direction, clusters categories by residual gradient compatibility, trains group-specific LoRA experts, and consolidates them into one adapter by weight merging followed by on-policy distillation from the experts. We motivate this procedure by viewing reward-weighted flow matching as velocity-field regression: incompatible reward dimensions may prefer different denoising directions in overlapping noisy latent regions, causing shared LoRA training to average capabilities. On Wan2.1-T2V-1.3B-Diffusers, the merged grouped adapter improves the matched VBench2.0 evaluation over the base model, a joint reward-weighted LoRA baseline, and random, semantic, and raw-gradient partitions trained with the same pipeline; an independent evaluator agrees, and on CogVideoX-2B grouping avoids the negative transfer of joint training. The gain is not uniform: merging compresses the largest specialist gains, distillation recovers part of this loss, and camera motion and several local-quality dimensions remain challenging. Together, these results suggest that gradient compatibility can serve as a practical diagnostic for organizing reward-weighted video post-training data.