Papers for

explainable ai developers

Papers whose findings have a practical use for this group, as judged from the abstract. Open a paper to read what it means in practice.

Vision sparse autoencoders challenge common interpretability measures

Verifying the Linear Representation Hypothesis: How Interpretable Are Vision SAEs?

Abstract: Vision Sparse Autoencoders (SAEs) have become a popular tool in Mechanistic Interpretability due to their presumed ability to disentangle complex features learned by a model into monosemantic concepts. Despite their growing popularity, evaluating their interpretability remains an active topic of research. The bedrock motivating the adoption of SAEs is the Linear Representation Hypothesis (LRH), which claims that polysemantic features can be projected onto a (near) orthogonal basis of sparse, human-understandable representations. Yet, most current frameworks evaluate proxies such as the sparsity of SAE features or the coherence of the inferred dictionary, implicitly assuming that these reflect alignment with human perception. In this paper, we provide empirical evidence that measuring the interpretability of SAE concepts is more difficult than these proxies suggest. To this end, we adapt the Autointerpretability Score (AIS) - previously shown to align with human judgments in Natural Language Processing - to vision tasks and validate our approach in a dedicated user study. We evaluate SAE concept quality using both standard metrics and our adapted AIS. We find that established interpretability metrics for SAEs correlate neither with one another nor with AIS, indicating that no single reference-free metric, whether grounded in the LRH or not, is sufficient for verifying the interpretability of vision SAEs. We argue these findings support recent calls for more verifiable, ground-truth-anchored design and evaluation of explanation methods.

Mon 28 SeptComputer Vision and Pattern Recognition
The gist
Understanding how artificial intelligence models see images is hard. Vision Sparse Autoencoders (SAEs) are tools designed to find simple, clear concepts inside complex image data, but it’s tricky to check how understandable these concepts really are. This paper finds that popular ways to judge these concepts don't agree with each other or with a new, user-tested method, meaning it’s harder than expected to verify how interpretable SAEs are. The authors suggest that more reliable and grounded ways to test AI explanations are needed.
Open → 2609.35020v1