Progressive training improves large multimodal models from crops to full images

Progressive-View On-Policy Distillation for Regional-to-Global Transfer in Multimodal LLMs

Computer Vision and Pattern RecognitionArtificial Intelligence

Summary

Understanding the whole picture in large multimodal models is hard when trained only on small parts or crops of images. The authors propose a step-by-step training method that gradually shifts a model’s focus from smaller image parts to the full image view while keeping important regional details. This method also adjusts the learning emphasis on parts of the image that the model finds challenging. Their approach improves the accuracy of models on several tasks involving images and questions compared to previous training techniques.

What this means in practice

Authors

Shanfeng Huang, Zhou Fang, Song Xiao, Hai Du

Abstract

Regional-to-global distillation uses crop-conditioned guidance to improve full-image understanding. The challenge is to effectively transfer the teacher's crop-based advantage to the student's full-image inference. We propose progressive-view on-policy distillation (PVD), which shifts the student's view distribution from the crop toward the full image through an intermediate aspect-preserving padded crop. The padded crop preserves regional content while matching the full image's visual-token grid. Across stages, the view mixture assigns increasing probability to the full image. A lightweight regional-advantage weighting reallocates token-level supervision using the crop-conditioned teacher-student log-probability gap. Evaluated under each sampled input, it applies mild reweighting when the gap is small and emphasizes higher-gap tokens when the gap widens. A Jensen-Shannon metric decomposition interprets this schedule as a shift from matched-input imitation toward the deployment objective. Across benchmarks spanning perception, visual mathematics and general multimodal question answering, PVD-full reaches an average accuracy of 77.51 over three seeds, improving on the reward-free distillation baseline by 2.01 points and on its reward-matched variant by 1.00 point. In the reward-free setting, PVD-distill still gains 1.16 points.