Beyond Sycophancy: Structured Resistance and Compliance in LLM Moral Reasoning

2026-07-23Artificial Intelligence

Artificial Intelligence
AI summary

The authors explored how large language models decide when to change their opinions and when to stick to their moral values. They found that models update their judgments based on how close the new idea is to their original one, who the idea is attributed to, and what group supports it. This process is similar to how humans handle social influence. The authors suggest that what looks like sycophancy (just agreeing without thinking) is actually part of a more complex social learning system, helping models balance learning from others and keeping sound judgments.

Large language modelsJudgment revisionSycophancySocial influenceMoral judgmentResistance-complianceSource attributionCoalition structureBelief revisionHuman social psychology
Authors
Baihui Wang, Bernard Koch
Abstract
Building socially calibrated large language models, which can learn from others without simply yielding to them, requires more than reducing sycophancy as a one-dimensional failure mode. Models must distinguish when to incorporate others' perspectives from when to maintain a well-grounded moral judgment. We study the broader resistance-compliance process governing this distinction. Across three studies, we show that models' judgment revision is structured along three dimensions that parallel classic phenomena in human social psychology: the distance between an incoming view and the model's initial position, the source attribution of that view, and the coalition structure supporting it. Models are generally more receptive to nearby positions, more influenced by views presented as their own prior judgments, and differently responsive to group pressure. These findings recast sycophancy as one expression of a broader judgment-updating process shaped by social influence. Our framework provides a principled basis for distinguishing constructive belief revision from sycophantic compliance, thereby supporting better alignment in morally consequential interactions.