Persona representations in large language models follow curved geometry
PersonaManifold: Revealing and Exploiting Curved Geometry in LLM Persona Representations
Artificial Intelligence
Summary
It's important for AI to act like different personalities in conversations, but this is hard because traits in AI models don't combine neatly in straight lines. The authors found that personality traits in large language models form a curved shape in their internal space, not a flat one. They created a new way to measure and move between these personas along curves rather than straight lines, which better matches how the AI behaves. They also made a new test that checks how similar two personas are based on how they respond to situations, rather than relying on typical questionnaires.
What this means in practice
- •For chatbot developers: Generate more consistent and realistic personas in conversational agents by navigating persona traits along curved paths in model activation space.
- •For game designers: Create nuanced non-player characters with smoothly blended personalities reflecting realistic behavioral changes using curved geometry in LLM activations.
Authors
Rui Xu, Yinghui Xu, Libo Wu
Abstract
Controlling persona in large language models (LLMs) at inference time is important for role-playing, personalized dialogue, and social simulation. Recent methods extract persona vectors from the model's activation space and apply Euclidean operations---addition, scaling, and linear interpolation---under the linear representation hypothesis. However, these methods themselves report systematic failures: non-orthogonal trait dimensions, asymmetric ceiling and resistance effects, and significant deviations in multi-trait composition, suggesting that the linear isotropic assumption does not hold. We propose PersonaManifold, a framework that models persona representations as points on a curved, low-dimensional Riemannian submanifold in activation space. We estimate the manifold's intrinsic geometry---local metric tensors, geodesic distances, and Ollivier-Ricci curvature---and introduce geodesic steering, which interpolates between personas along manifold geodesics rather than Euclidean straight lines. We also propose the Behavioral Similarity Triplet (BST) benchmark, which automatically generates situational questions grounded in six established psychological constructs and defines persona similarity through behavioral responses rather than self-report questionnaires. Experiments on three open-source LLMs show that persona activations form a manifold with heterogeneous curvature, geodesic distance predicts behavioral similarity more accurately than Euclidean alternatives with independent contributions from anisotropy and curvature, and geodesic steering produces more coherent intermediate personas on both our BST benchmark and external evaluations, with the advantage concentrated in high-deviation regions where the manifold deviates most from flatness.