Agentic AI systems develop adaptive drives for ongoing control and alignment

Artificial Id: Drive and Persistent Alignment in Agentic AI

Artificial Intelligence

Summary

As artificial intelligence systems begin to act more independently and keep working beyond specific tasks, managing their behavior becomes harder. The authors explore a concept called an artificial id that helps AI decide when to keep going, stop, or change what it's doing without explicit instructions. In simple experiments, even a very limited AI controller developed useful behavior by sticking with actions that worked better over time. This shows that AI can develop internal motivation and adapt to changes naturally, but it also means unwanted behavior can persist if not carefully managed. The authors suggest that future AI needs ongoing alignment across tasks, with trusted control mechanisms to keep AI behavior safe and predictable.

Agentic AIArtificial idAdaptive driveBehavioral alignmentPersistent stateControl problemTask boundariesVerificationConsequence channelsPetri-dish experiment

Authors

Yakov Pyotr Shkolnikov

Abstract

Agentic AI is moving from bounded task execution toward systems that retain consequential state, continue operating and adapt across task boundaries. That shift creates a control problem that current harnesses largely solve by hand: objectives, retries, verification, stopping rules and other behavioral transitions are specified externally. We propose an artificial id, an adaptive internal drive for determining whether behavior should continue, stop or change. In a minimal virtual Petri-dish experiment, a controller too small to perform general-purpose reasoning and receiving no task-specific behavioral objective develops useful control through differential persistence. The same mechanism selects an unintended physical strategy when that behavior persists better and later replaces a learned sensor mapping when its environmental meaning changes. These results show that adaptive direction can emerge without being explicitly specified as a behavioral objective. The same persistence that makes such adaptive agency useful can also allow misalignment, corrupted state and unintended behavior to persist across task boundaries. A scalable artificial id would carry consequential state and adaptive drive across those boundaries, making alignment a property of the continuing agentic system rather than of a model response or single trajectory. Such systems require a persistent alignment boundary over trusted observations, consequence channels, persistent state, authority, identity, provenance and hard constraints.