Cosine similarity can miss key dialogue model structures despite linear access

When Cosine Similarity Fails to Reflect Linearly Accessible Structure in Dialogue Models

Computation and Language

Summary

Cosine similarity is a common way to measure how alike two sets of words or meanings are inside language models. The authors show that in chat-focused language models, cosine similarity often doesn’t capture important personality-related patterns that are actually easy to find with other methods. They find that some ways to focus on important features recover these patterns, but simple dimension reduction like PCA doesn’t help. This problem doesn’t happen in simple tasks like sentiment analysis, and it doesn’t get worse over the course of a conversation.

What this means in practice

  • For language model developers: Improve evaluation of dialogue models by using linear probes instead of relying solely on cosine similarity for capturing persona features.
  • For ai testing teams: Design tests for chatbots that check for task-relevant structure beyond ambient similarity measures to better assess model understanding.

Authors

Yu Sun, Mengyin Lu, Cong Feng, Guangming Lu, Huimin Han

Abstract

Cosine similarity is widely used to analyze transformer representations, implicitly assuming that similarity reflects task-relevant structure. We study when this assumption fails in dialogue-conditioned large language models. Across three 7-8B chat-tuned models, ambient cosine similarity substantially underestimates linearly decodable persona structure on the same hidden states; numerically, linear probe AUC is in the 0.73-0.97 range while cosine kNN is in the 0.56-0.77 range on a 30-class task. A low-dimensional supervised subspace recovers much of this gap, whereas a matched-rank PCA subspace does not and in some cases degrades performance. This mismatch is regime-dependent: it is absent in single-sentence sentiment classification (SST-5), and a matched-cardinality control rules out attribute cardinality as a confound. The gap does not systematically increase across dialogue turns, and the task-aligned subspace remains stable over time. However, two of three models violate a pre-registered within-subspace separability invariance criterion (|Delta AUC| <= 0.03), and one model violates a pre-registered turn-invariance criterion (|Delta L| <= 0.05). These results show that cosine similarity can fail to reflect task-aligned structure in dialogue representations even when that structure is linearly accessible.