Multimodal transformers build virtual encoders inside their layers
Virtual Encoders in Multimodal Transformers
Computer Vision and Pattern Recognition
Summary
Multimodal language models usually use separate parts to understand images, sounds, or other sensory data before using language. The authors found that some advanced models can do this sensory processing inside their own transformer layers without separate encoders. These internal computations, called Virtual Encoders, happen early inside the model and make the sensory data ready for language tasks. This changes how we think about where perception and language processing happen in multimodal AI.
What this means in practice
- •For natural language processing engineers: Design multimodal models without dedicated perceptual encoders by relying on transformers to internally process sensory data.
- •For multimedia software developers: Create efficient multimodal AI applications by embedding encoding computation within a shared transformer, reducing architecture complexity.
Authors
Katsuya Ogata, Yuta Nakashima
Abstract
Multimodal language models traditionally rely on dedicated perceptual encoders to construct task-usable representations. More integrated architectures have recently emerged, which instead expose the shared transformer to lightly projected patches, audio frames, or discrete visual tokens. Where does this encoding happen when such representations are not provided? We find that the transformer can internalize this missing computation, constructing task-usable perceptual representations within its own early-to-middle layers before the downstream language model. We call this computational structure a Virtual Encoder. Across linear probing, similarities to perceptual encoders, and causal analyses, we identify signatures of this structure in models that receive perceptual tokens without continuous encoder-derived features. These analyses also suggest that the boundary between perception and language processing need not coincide within an architectural module. Instead, encoder-like computation can emerge as a functional regime within a shared transformer, providing a new perspective for understanding where and how multimodal models process perception.