How Language Models Organize and Structure Moral Knowledge
2026-08-27 • Computation and Language
Computation and LanguageArtificial IntelligenceMachine Learning
AI summaryⓘ
The authors studied how large language models understand different types of moral ideas. They trained six probes to detect specific moral categories and found that the models do not treat all morals the same way but organize them into related, distinct patterns with some shared features. This organization appears early during training and is consistent across different models, reflecting how morals appear in the data rather than human moral theory divisions. When analyzing moral dilemmas, the models capture both the combination of moral foundations and the unique conflicts in each dilemma, showing they represent moral tension rather than fixed judgments.
Large Language ModelsMoral Foundations TheoryLinear ProbesRepresentation SpaceMoral CategorizationEmbedding GeometryMoral DilemmasModel Pre-trainingMoral TensionCorpus Statistics
Authors
Orion Reblitz-Richardson
Abstract
How do large language models (LLMs) organize moral knowledge? Models detect moral content broadly, but detection is a low bar. We ask whether they go further, distinguishing moral foundations from one another and organizing the relationships between them geometrically. We train six independent linear probes on open-weight language models, one per Moral Foundations Theory (MFT) category (care/harm, fair/cheat, lib/oppress, loy/betray, auth/subv, sanc/degrade), and examine how the resulting directions relate to each other in representation space. We find the directions neither collapse into a single moral detector nor isolate from one another. Rather, they span a near-maximal number of independent dimensions while sharing a positive common component. The shared component is the signature of integration, and it is moral-specific relative to a matched non-moral concept battery built identically (mean pairwise cosine 0.26 vs. 0.013). The geometry is consistent across architectures and scale and reaches its integration regime early in pre-training, well before probe accuracy saturates. The structure the model discovers shows no evidence of the individualizing/binding distinction predicted by Moral Foundations Theory (an underpowered test: only 20 candidate partitions exist) but rather reflects corpus statistics. Extending to moral dilemmas, each dilemma direction partially composes from its component foundations, at 2.7x a mismatched-pair baseline, while the majority of its variance encodes conflict-specific structure. The model represents moral tension itself, not a pre-resolved judgment.