Papers for

automated visual question answering developers

Papers whose findings have a practical use for this group, as judged from the abstract. Open a paper to read what it means in practice.

Progressive training improves large multimodal models from crops to full images

Progressive-View On-Policy Distillation for Regional-to-Global Transfer in Multimodal LLMs

Abstract: Regional-to-global distillation uses crop-conditioned guidance to improve full-image understanding. The challenge is to effectively transfer the teacher's crop-based advantage to the student's full-image inference. We propose progressive-view on-policy distillation (PVD), which shifts the student's view distribution from the crop toward the full image through an intermediate aspect-preserving padded crop. The padded crop preserves regional content while matching the full image's visual-token grid. Across stages, the view mixture assigns increasing probability to the full image. A lightweight regional-advantage weighting reallocates token-level supervision using the crop-conditioned teacher-student log-probability gap. Evaluated under each sampled input, it applies mild reweighting when the gap is small and emphasizes higher-gap tokens when the gap widens. A Jensen-Shannon metric decomposition interprets this schedule as a shift from matched-input imitation toward the deployment objective. Across benchmarks spanning perception, visual mathematics and general multimodal question answering, PVD-full reaches an average accuracy of 77.51 over three seeds, improving on the reward-free distillation baseline by 2.01 points and on its reward-matched variant by 1.00 point. In the reward-free setting, PVD-distill still gains 1.16 points.

Sat 26 SeptComputer Vision and Pattern RecognitionArtificial Intelligence
The gist
Understanding the whole picture in large multimodal models is hard when trained only on small parts or crops of images. The authors propose a step-by-step training method that gradually shifts a model’s focus from smaller image parts to the full image view while keeping important regional details. This method also adjusts the learning emphasis on parts of the image that the model finds challenging. Their approach improves the accuracy of models on several tasks involving images and questions compared to previous training techniques.
Open → 2609.32333v1