Sparse autoencoders reveal parts of speech in language data
Parts-of-Speech as Emergent Categories in SAE Latent Space
Computation and Language
Summary
The paper studies how a type of AI model called a Sparse AutoEncoder (SAE) represents parts of speech like verbs and nouns. The authors find that these models can identify parts of speech well, but not by matching one part-of-speech category to one specific signal inside the model. Instead, groups of signals together represent these categories. The system organizes language information in a complex, spread-out way rather than using simple, single features.
What this means in practice
- •For natural language processing engineers: Design better language analysis tools by exploiting distributed part-of-speech signals in sparse autoencoder outputs.
- •For computational linguists: Use groupings of latent features to study morpho-syntactic categories rather than relying on isolated model units.
Authors
Alessandro Bondielli, Lucia Passaro, Serena Auriemma, Alessandro Lenci
Abstract
Sparse AutoEncoders (SAEs) offer a promising way to inspect language model representations, but it is still unclear what kind of linguistic structure their latents expose. We use part-of-speech (PoS) categories as a controlled test case to study whether morpho-syntactic information is encoded by individual latents or by structured groups of features. We find that PoS distinctions are highly recoverable from SAE activations, but do not align with one-to-one latent / category mappings. This recoverability is not reducible to lexical memorisation, and Open and Closed PoS classes differ substantially. Categories are supported by compact groups of sparse latents, with substantial variation across tags. These groups remain stable on held-out data, while also showing overlap between related categories. Our results show that SAEs localise morpho-syntactic information in a distributed and category-dependent form rather than through atomic grammatical features.