Self emergence architecture creates distinct stable agent personalities
Self-Emergence Agent Architecture:Behavior-Inertia HMM, Reflexive Metacognition,and Social-Contrastive Self-Modeling
Artificial Intelligence
Summary
Language model agents can generate language well but often struggle with keeping a consistent personality and recognizing themselves versus others. The authors created a new system where agents update their behavior based on past actions, think about their own thoughts, and compare themselves with other agents. This loop helps initially identical agents develop stable, different personalities and social roles. Their experiments show this method forms clear social structures and distinct self-narratives that previous designs do not produce.
What this means in practice
- •For ai system developers: Design agent behaviors that evolve distinct personalities and social roles through continuous self-reflection and social comparison loops.
- •For interactive game developers: Create more believable non-player characters that develop unique, stable behavior patterns and social dynamics over time.
Authors
Xiaoyang Liu
Abstract
Large language model (LLM) agents exhibit strong language-generation and problem-solving capabilities, yet suffer from three structural limitations: personality drift, non-evolutionary reflection, and the absence of a self-other boundary. Existing generative-agent simulations rely on static memory and fixed prompts, maintaining neither behavioral inertia nor endogenous self-evolution. We propose the Self-Emergence Agent Architecture (SEAA), which integrates three components: (i) a Hidden Markov Model (HMM) that encodes long-term behavioral and cognitive inertia as an editable state-transition matrix; (ii) a Reflexion-style verbal metacognition loop whose output updates the HMM parameters themselves, rather than merely being stored as text; and (iii) a multi-agent social environment in which initially identical agents continuously compare their behavior with others'. The three components form a closed loop: social action $\to$ feedback $\to$ self-reflection $\to$ inertia update $\to$ differentiated action. We state three falsifiable hypotheses and provide a reproducible experimental protocol with operational metrics. A language-model-free prototype shows the loop spontaneously breaks symmetry: initially identical agents consolidate distinct, stable personalities whereas matched controls do not. Experiments with a hosted LLM surface these differences as distinct first-person self-narratives, and a five-agent deliberation spontaneously develops social structure---a consensus hub and a unanimously rejected outlier---absent in the control. Following an epistemologically agnostic stance inspired by Zhuangzi, SEAA studies only observable behavioral emergence and makes no claim about subjective qualia. This work contributes a unified framework, a concrete architecture with pseudocode, mechanistic evidence, and a microscope-style sandbox for studying artificial-self emergence.