Pretrained convolutional features boost vision language models accuracy
ConvCue: Complementary Visual Inductive Biases for Vision-Language Models
Computer Vision and Pattern Recognition
Summary
Vision-language models can understand pictures and text together, but sometimes they struggle with detailed visual tasks like telling apart similar objects or understanding spatial layouts. The authors showed that adding extra visual information from a special type of image analyzer called a convolutional neural network (CNN) can help these models do better. They kept the original model’s visual part and added the CNN features alongside it, letting the model learn to combine both kinds of information. This approach improved performance on many different visual and language tasks without replacing the original tools.
What this means in practice
- •For multimodal ai engineers: Improve model accuracy on visual question answering and reasoning tasks by adding pretrained convolutional features alongside existing visual encoders.
- •For document processing teams: Enhance automated understanding of documents and charts by integrating convolutional features to better capture spatial and fine visual details.
Authors
Zixuan Lan, Shichu Sun
Abstract
Modern vision-language models (VLMs) achieve strong performance across a broad range of multimodal tasks, yet still struggle with visual questions that require fine-grained discrimination and spatial understanding. These limitations motivate investigating whether supplementary visual representations can improve existing VLMs without replacing their native visual encoders. Pretrained convolutional networks offer a candidate feature source, motivated by their local connectivity and spatial weight sharing. We introduce CONVCUE, which augments the native visual representations of a pretrained VLM with final-stage features from a parallel, frozen pretrained CNN. A learnable adapter maps convolutional features to the native visual feature dimension, while gated cross-attention allows the original visual tokens to retrieve information from the CNN features. The enhanced tokens are passed through the original visual-to-language projector, and the model is adapted through a two-stage training procedure. We evaluate CONVCUE on Qwen3-VL-2B, Qwen3-VL-4B, and LLaVA-OneVision-7B across 13 multimodal benchmarks covering visual question answering, document and chart understanding, and multimodal reasoning. CONVCUE improves average benchmark performance over both the original models and matched two-stage fine-tuning controls on all three backbones. On Qwen3-VL-4B, it improves over the original model on all 13 benchmarks and raises the average score from 75.00 to 78.82 relative to the matched fine-tuning control. These results show that pretrained convolutional representations, when integrated through learned adaptation and fusion, can improve the visual understanding of existing VLMs without replacing their original visual encoders.