Papers for

multimodal ai engineers

Papers whose findings have a practical use for this group, as judged from the abstract. Open a paper to read what it means in practice.

Gradient-guided grouping improves reward-weighted video model training

G$^3$-LoRA: Organizing Reward-Weighted Video Data with Gradient-Guided Grouped LoRA

Abstract: Post-training foundation video models on heterogeneous reward-weighted data usually assume that all data categories induce compatible updates. This assumption is fragile when categories correspond to different skills, domains, or evaluation dimensions. We study this problem in text-to-video post-training, where VBench2.0 dimensions define data buckets and an external multimodal reward pipeline assigns sample weights. We propose G$^3$-LoRA (Gradient-Guided Grouped LoRA), a data organization procedure that probes category-level gradients induced by reward-weighted video samples, removes the shared global update direction, clusters categories by residual gradient compatibility, trains group-specific LoRA experts, and consolidates them into one adapter by weight merging followed by on-policy distillation from the experts. We motivate this procedure by viewing reward-weighted flow matching as velocity-field regression: incompatible reward dimensions may prefer different denoising directions in overlapping noisy latent regions, causing shared LoRA training to average capabilities. On Wan2.1-T2V-1.3B-Diffusers, the merged grouped adapter improves the matched VBench2.0 evaluation over the base model, a joint reward-weighted LoRA baseline, and random, semantic, and raw-gradient partitions trained with the same pipeline; an independent evaluator agrees, and on CogVideoX-2B grouping avoids the negative transfer of joint training. The gain is not uniform: merging compresses the largest specialist gains, distillation recovers part of this loss, and camera motion and several local-quality dimensions remain challenging. Together, these results suggest that gradient compatibility can serve as a practical diagnostic for organizing reward-weighted video post-training data.

Mon 28 SeptComputer Vision and Pattern RecognitionMachine Learning
The gist
Training video models using various types of reward signals can be hard because different types of feedback might suggest conflicting improvements. The authors study a way to organize training data by grouping similar types of feedback according to how their training updates interact. They propose a new method called G3-LoRA that clusters these groups based on gradient directions, trains specialized adapters for each, and then merges them to improve the overall model. This method leads to better performance compared to mixing all feedback together, though some specialized skills remain challenging.
Open → 2609.35189v1

Pretrained convolutional features boost vision language models accuracy

ConvCue: Complementary Visual Inductive Biases for Vision-Language Models

Abstract: Modern vision-language models (VLMs) achieve strong performance across a broad range of multimodal tasks, yet still struggle with visual questions that require fine-grained discrimination and spatial understanding. These limitations motivate investigating whether supplementary visual representations can improve existing VLMs without replacing their native visual encoders. Pretrained convolutional networks offer a candidate feature source, motivated by their local connectivity and spatial weight sharing. We introduce CONVCUE, which augments the native visual representations of a pretrained VLM with final-stage features from a parallel, frozen pretrained CNN. A learnable adapter maps convolutional features to the native visual feature dimension, while gated cross-attention allows the original visual tokens to retrieve information from the CNN features. The enhanced tokens are passed through the original visual-to-language projector, and the model is adapted through a two-stage training procedure. We evaluate CONVCUE on Qwen3-VL-2B, Qwen3-VL-4B, and LLaVA-OneVision-7B across 13 multimodal benchmarks covering visual question answering, document and chart understanding, and multimodal reasoning. CONVCUE improves average benchmark performance over both the original models and matched two-stage fine-tuning controls on all three backbones. On Qwen3-VL-4B, it improves over the original model on all 13 benchmarks and raises the average score from 75.00 to 78.82 relative to the matched fine-tuning control. These results show that pretrained convolutional representations, when integrated through learned adaptation and fusion, can improve the visual understanding of existing VLMs without replacing their original visual encoders.

Mon 28 SeptComputer Vision and Pattern Recognition
The gist
Vision-language models can understand pictures and text together, but sometimes they struggle with detailed visual tasks like telling apart similar objects or understanding spatial layouts. The authors showed that adding extra visual information from a special type of image analyzer called a convolutional neural network (CNN) can help these models do better. They kept the original model’s visual part and added the CNN features alongside it, letting the model learn to combine both kinds of information. This approach improved performance on many different visual and language tasks without replacing the original tools.
Open → 2609.34196v1

Attention methods keep results stable despite signal changes in transformers

Refinement Symmetry in Multimodal Transformers

Abstract: Attention weights depend on token counts, which change with the representation of a signal. We study refinement symmetry: splitting a representation while preserving content, position, visible context, and total mass should preserve its contribution. Building on proportional and quadrature attention, we show that split invariance forces the local mass factor to be linear for any fixed positive attention kernel, provided that factor is nondecreasing. For changed representations, a physical coupling bounds attention error by separating feature change from weight reallocation. In Qwen2.5-Omni-7B, duplicating half the visual tokens threefold changes 255 of 3,586 MVBench answers under standard attention; measure weighting preserves every answer under matched visibility. Under natural frame resampling, it reduces distributional drift. At twofold merging of a frozen video encoding, a five-seed evaluation shows an all-partition-correct accuracy gain of 1.04 percentage points over global count weighting (average group mass) and 0.93 points over standard attention. The advantage over global count also holds on WorldSense but depends on the compression budget. The result is a representation principle with a measured benefit in robustness across partitions.

Sat 26 SeptArtificial IntelligenceComputer Vision and Pattern Recognition
The gist
Attention in multimodal transformers depends on how many pieces (tokens) a signal is split into, which can change its influence unexpectedly. The authors studied how breaking signals into smaller parts while keeping content and position the same should not change results. They developed a method that weighs attention to keep answers consistent even when the tokens representing images or videos are duplicated or merged differently. Their approach improves stability and accuracy in real models handling video and multimodal data.
Open → 2609.32669v1

Progressive training improves large multimodal models from crops to full images

Progressive-View On-Policy Distillation for Regional-to-Global Transfer in Multimodal LLMs

Abstract: Regional-to-global distillation uses crop-conditioned guidance to improve full-image understanding. The challenge is to effectively transfer the teacher's crop-based advantage to the student's full-image inference. We propose progressive-view on-policy distillation (PVD), which shifts the student's view distribution from the crop toward the full image through an intermediate aspect-preserving padded crop. The padded crop preserves regional content while matching the full image's visual-token grid. Across stages, the view mixture assigns increasing probability to the full image. A lightweight regional-advantage weighting reallocates token-level supervision using the crop-conditioned teacher-student log-probability gap. Evaluated under each sampled input, it applies mild reweighting when the gap is small and emphasizes higher-gap tokens when the gap widens. A Jensen-Shannon metric decomposition interprets this schedule as a shift from matched-input imitation toward the deployment objective. Across benchmarks spanning perception, visual mathematics and general multimodal question answering, PVD-full reaches an average accuracy of 77.51 over three seeds, improving on the reward-free distillation baseline by 2.01 points and on its reward-matched variant by 1.00 point. In the reward-free setting, PVD-distill still gains 1.16 points.

Sat 26 SeptComputer Vision and Pattern RecognitionArtificial Intelligence
The gist
Understanding the whole picture in large multimodal models is hard when trained only on small parts or crops of images. The authors propose a step-by-step training method that gradually shifts a model’s focus from smaller image parts to the full image view while keeping important regional details. This method also adjusts the learning emphasis on parts of the image that the model finds challenging. Their approach improves the accuracy of models on several tasks involving images and questions compared to previous training techniques.
Open → 2609.32333v1

Vision language models assessed for true visual causal reasoning skills

CCRV-Bench: Constraint-Based Evaluation of Causal Reasoning in Vision-Language Models

Abstract: Vision-language models (VLMs) have demonstrated excellent performance in visual tasks, but their visual causal reasoning capabilities still lack reliable evaluation. Existing evaluations struggle to distinguish whether a model is performing causal reasoning based on visual evidence or relying on statistical correlations for shortcut learning, thereby potentially overestimating their actual capabilities. This paper proposes CCRV-Bench, a constraint-driven visual causal reasoning benchmark for single-image physical scenarios. We construct an orthogonal framework that evaluates four causal task dimensions: causal relation discovery, state prediction, causal diagnosis, and intervention. We further introduce entity symbolization, spatial grounding, the factual adversarial constraint, and minimalist output constraints to reduce shortcut cues while preserving the physical commonsense required by the task. Experiments across 15 multimodal models show that constraint sensitivity is task- and model-dependent: intervention has the largest average effective degradation among the four causal tasks, spatial grounding is the most damaging constraint on average, and the factual adversarial constraint improves DCR for all evaluated models. These results show that unconstrained performance does not determine constrained robustness and that a single aggregate score can obscure distinct failures in causal identification, spatial grounding, and constraint-compliant expression. CCRV-Bench provides a standardized framework for diagnosing image-grounded causal reasoning under controlled constraints. The code is available at https://github.com/0815linyuan/CCRV-Bench-Constraint-Based-Evaluation-of-Causal-Reasoning-in-Vision-Language-Models

Fri 25 SeptComputer Vision and Pattern Recognition
The gist
Many AI systems that combine seeing and understanding language appear to do well on tasks, but it’s unclear if they really understand cause and effect from images or just guess based on patterns they’ve seen before. The authors created a special test called CCRV-Bench to better check if these models truly reason about causes in single images by controlling tricky clues that can mislead AI. They tested 15 models and found that some tasks and constraints reveal weaknesses that overall scores hide, showing that good scores don’t always mean real understanding. Their benchmark helps pinpoint exactly where these models succeed or fail in causal thinking grounded in what they see.
Open → 2609.30979v1

Speech content features compared on audio generation and identity control

A Comprehensive Study of Content Representations for Speech Synthesis

Abstract: Speech content representations are central to voice conversion, speech-to-speech translation, and multimodal language models, yet they are rarely compared under a common generative framework that directly measures what each representation contains. We address this by training a generative model conditioned solely on each representation and evaluating the generated audio along the content, speaker identity, and prosody axes. Across SSL features, supervised tokens, posteriorgrams, and neural audio codecs, we find two distinct regimes: representations that nearly reconstruct the original audio, and representations that effectively disentangle speaker identity. These results show that disentanglement depends not on supervision alone, but on the interaction between the training objective and the representation's information capacity: supervised representations only disentangle speaker identity when their capacity is sufficiently constrained.

Fri 25 SeptMachine LearningSound
The gist
It can be hard to tell how different ways of representing spoken language actually affect the sounds that a computer makes when it tries to copy or change someone's voice. This study trained a computer model to recreate speech using different types of speech representations, then checked how well each kept the words, the speaker’s unique voice, and the rhythm of speech. The authors found that some representations do a great job of copying the original sound, while others better separate who is speaking from what is being said. They showed that whether or not the speaker’s identity can be separated depends not just on the method used, but also on how much information the representation can hold.
Open → 2609.30975v1

Layer aware position embeddings improve visual token pruning in multimodal models

Layer-Aware Position Embeddings for Visual Token Pruning in Multimodal Large Language Models

Abstract: Multimodal large language models (MLLMs) incur substantial computational overhead due to the reliance on hundreds of visual tokens to represent images. While token pruning has emerged as a promising approach to reduce the inference cost of MLLMs, existing methods typically reassign position embeddings to the retained tokens using either sparse or continuous position embeddings, each introducing distinct limitations. Sparse position embeddings tend to decrease the attention value allocated to visual tokens, thereby degrading the perception capability of MLLMs, whereas continuous position embeddings disrupt the original spatial correspondence of visual tokens, leading to weakened grounding capability. To mitigate this issue, we perform layer-wise analysis of the language decoder and observe that intermediate layers play a critical role for maintaining the grounding capability of MLLMs under token pruning. Based on this observation, we propose a layer-aware position embedding strategy, which switches to sparse position embeddings at grounding-sensitive layers while maintaining continuous position embeddings elsewhere. Extensive experiments across representative pruning methods and diverse benchmarks demonstrate that our approach improves the comprehensive multimodal performance of pruned MLLMs compared with standard sparse and continuous position embeddings.

Sun 20 SeptComputer Vision and Pattern Recognition
The gist
Multimodal large language models process images by breaking them into many small pieces called visual tokens, which takes a lot of computing power. To speed things up, some tokens can be removed, but this can confuse the model because the positions of the remaining tokens get reassigned in different ways, each with drawbacks. The authors found that some parts of the model care more about keeping tokens in their original positions to understand images well. They designed a method that changes how positions are reassigned depending on the layer, improving overall image understanding while making the model faster.
Open → 2609.23715v1