Visual encoders become less crucial for large multimodal models
How Far Are We from Removing the Visual Encoder? Scaling Laws for Encoder-Free Multimodal Pretraining
Computer Vision and Pattern Recognition
Summary
Many systems that understand images and text use a special visual encoder to help interpret pictures. This paper shows that if you remove this visual encoder, the model can still learn to understand images, but it needs to be much bigger and use more computing power. The researchers found that small models without this encoder don't do as well, but very large models can catch up and learn to handle images internally. This means that for very big AI models, we might no longer need separate visual encoders, simplifying the design.
What this means in practice
- •For machine learning engineers: Design larger multimodal AI systems that can learn visual features directly from raw images without relying on pretrained visual encoders.
- •For software developers for ai platforms: Simplify multimodal AI architectures by removing separate visual encoders for improved model integration and maintenance at large scales.
Authors
Lin Chen, Bolin Ni, Qi Yang, Lan Jiang, Kun Ding, Xiaoran Fan, Hower Yang, Ying Wang, Shiming Xiang
Abstract
Most modern multimodal large language models (MLLMs) build on a pretrained visual encoder that provides a strong visual prior. Encoder-free MLLMs instead learn visual representations directly from raw pixels, offering a simple and unified architecture, but their scaling behavior has not been systematically characterized. To fill this gap, we compare scaling laws for encoder-free and encoder-based MLLMs and report three main findings: (1) Removing the visual encoder shifts the compute-optimal allocation for the multimodal objective toward larger models, while leaving that for text nearly unchanged. (2) The two architectures exhibit nearly overlapping loss--compute frontiers on the text objective, but diverge on the multimodal objective: encoder-free models underperform at small scales yet are predicted to catch up at around $10^{22}$ FLOPs, well within practical pretraining budgets. (3) Without a visual encoder, the language model learns to take over its role via vision-specific adaptation: bidirectional interactions among visual tokens become increasingly beneficial as training compute grows, visual processing shifts toward earlier layers, and expert routing for visual tokens becomes more concentrated. Overall, our results indicate that the advantage of the visual prior provided by a pretrained encoder diminishes with scale, positioning encoder-free architectures as a promising direction for multimodal pretraining.