When Structured Sparse Autoencoders Learn Consistent Concepts Across Modalities

2026-07-09Computer Vision and Pattern Recognition

Computer Vision and Pattern RecognitionArtificial IntelligenceMachine Learning
AI summary

The authors study how to improve sparse autoencoders (SAEs) used in vision-language models, which try to learn distinct concepts from images and text. They find that regular SAEs often struggle to keep visual concepts consistent and connected in images. To fix this, they create a Structured Sparse AutoEncoder (S2AE) that groups image patches by how related they are in attention and location, encouraging latent features to represent clearer and more consistent concepts. Their method improves the alignment of learned concepts with actual image parts, makes the representation more efficient, and helps the model produce cleaner, more distinct features in both vision and language.

Sparse AutoencoderVision-Language ModelsConcept ConsistencyTransformer AttentionStructured SparsitySemantic AlignmentLatent FeaturesMonosemanticityReconstruction Fidelity
Authors
Weiduo Liao, Yunqiao Yang, Ying Wei
Abstract
Sparse autoencoders (SAEs) have emerged as a promising technique for mechanistic interpretability by learning a set of sparse latent features in large models, each of which encodes a distinct concept. However, in vision-language models (VLMs), vanilla SAEs struggle to learn modality-consistent concepts, with concepts often exhibiting fragmented coverage (i.e., disjoint regions) in the visual modality. To address this challenge, we propose a Structured Sparse AutoEncoder ($S^2AE$) that enforces concept consistency from both semantic and spatial perspectives in the visual modality. Specifically, we group image patches based on Transformer attention similarity and spatial proximity, and introduce a structured sparsity regularization when training the vanilla SAE. The regularization consists of exclusive sparsity for inter-group concept disentanglement and group sparsity for intra-group concept consistency, which drives the latent neurons by SAEs to specialize in distinct, semantically grounded concepts. Evaluated on the \texttt{Qwen2.5-VL-7B-Instruct} model, the method achieves 6.06% average improvement in semantic alignment (mIoU) and 60.81 in representational efficiency (lower l0 norm) while maintaining near-perfect reconstruction fidelity with an Explained Variance above 99%. Cross-modal analysis further demonstrates that $S^2AE$ enhances neuronal monosemanticity by this visual structural prior, achieving a 3.08% average gain in semantic consistency and a 2.37% average gain in monosemanticity scores for both modalities of multimodal features, thereby fostering more coherent and disentangled representations.