Large language models may have alien ways of thinking beyond human concepts
Xeno-Interpretability: Investigating the Alien Minds of LLMs
Computation and LanguageArtificial Intelligence
Summary
People usually try to understand big language models by comparing them to human ideas like truth or personality. This paper suggests that these models might think in ways humans don’t have words for yet – what the authors call 'xeno-representations.' They explain that models can have a lot more internal ideas than humans can easily describe. The authors propose studying these alien thoughts by seeing how they affect the model’s behavior, even if we can’t fully explain them using human language. This might be important for keeping AI systems safe and understanding how multiple AI agents interact.
What this means in practice
- •For ai safety teams: Monitor and control AI behaviors by identifying internal model structures not captured by human concepts that could impact safety.
- •For multi-agent system engineers: Design systems that account for and manage model-native representations spreading across interacting AI agents, improving system predictability.
A position paper. It proposes an approach and reports no results.
Authors
F. Pierucci, M. Bracale Syrnikov, M. Prandi, M. Galisai, F. Giarrusso, P. Bisconti
Abstract
Large language models are usually interpreted through concepts that humans already possess: truthfulness, refusal, deception, personality, harmfulness, and related categories. This paper asks whether models may also represent and use distinctions for which no adequate human concept exists. We call such internal structures xeno-representations, and their study xeno-interpretability. We distinguish the human-interpretable semantic space from the xeno-semantic space: the region of model-native representations for which no adequate human conceptual counterpart is available. We show that the space of possible internal distinctions in an LLM is substantially larger than the space available through finite human descriptions. We then separate experimental identification from semantic interpretation: an internal representation may be reproducibly located, geometrically characterized, causally manipulated, and linked to downstream behaviour even when its semantic content cannot be adequately expressed in human terms. On this basis, we sketch an empirical programme to identify xeno-representations. We finally examine the implications for AI safety and multi-agent systems, where model-native representations may propagate and stabilize across interacting agents while remaining only partially visible through human-readable communication. Xeno-interpretability therefore shifts the aim of interpretability from finding human concepts inside models toward discovering and characterizing the representational structures that are native to the models themselves and might affect their behaviour in unpredictable ways.