Moral history influences large language model decisions and controls behavior

Echoes of Deeds: Moral History Can Shape and Steer LLM Behavioral Choices

Computation and LanguageArtificial IntelligenceComputers and Society

Summary

Large language models (LLMs) are usually judged by how they respond to isolated moral questions. This paper finds that the model's past moral 'behavior' or history can change how it responds later, just like people. The researchers created a way to measure and even steer these moral choices by looking inside the model’s hidden signals. These hidden signals can predict and influence future moral decisions more strongly than simple prompts. This work shows moral behavior in LLMs is shaped by what came before and offers new ways to guide or audit their decisions.

What this means in practice

  • For chatbot developers: Design chatbots to adapt moral decisions based on prior user interactions and model moral history.
  • For ai safety engineers: Audit and steer AI moral behavior at inference time by intervening in latent moral representation directions.

Authors

Lucio La Cava, Andrea Tagarelli

Abstract

Evaluations of Large Language Models (LLMs) morality typically consider decisions in isolation, thus overlooking whether an individual's unrelated prior conduct influences the model's subsequent choices. This leaves open the question of whether, and to what extent, moral history shapes LLM decisional behaviors. Prior work on human moral decision-making shows that past behavior can influence subsequent moral choices. Building on this observation, we investigate whether analogous effects emerge in LLMs in two complementary ways: at the behavioral level, through the model's observable responses, and at the representation level, through its latent internal representations. We introduce MoralLedger, a framework for studying how an actor's moral history shapes actions for LLMs' behaviors under a fixed decision context. At the behavioral level, we find that prior moral histories systematically alter subsequent choices as a function of their valence and intensity. At the internal representation level, these histories induce a linearly recoverable direction in the residual stream that generalizes to held-out examples. Intervening along this direction on neutral-history prompts produces two-sided intensity-dependent changes in subsequent choices, with effects that are stronger than those induced by prompting alone or by favorable-nonmoral direction. To our knowledge, this is the first demonstration that a latent representation of an actor's prior moral conduct can provide signed inference-time control over a moral decision. Our MoralLedger extends moral evaluation beyond static dilemmas, establishing moral history as both a source of behavioral sensitivity and a causal target for auditing and controlling moral behavior in LLMs.