How source and argument together shape ai evaluation judgments
The Argument and the Letterhead: Source-Position Coherence in AI Evaluation
Artificial Intelligence
Summary
People’s judgments about arguments change depending on who is said to have made them, even if the argument itself is the same. The authors studied this by asking many evaluators to rate fixed arguments that were presented as coming from different sources, like political groups or countries. They found that people’s ratings vary in a way that suggests they expect certain positions to fit certain sources, not just based on argument quality alone. This shows that AI evaluation needs to consider how the attributed source affects the evaluation of the argument.
What this means in practice
- •For ai system builders: Design evaluation methods that account for how source attributions influence argument ratings to improve interpretability of AI assessments.
- •For policy analysts: Better understand how biases tied to source and position may shape public and expert evaluations of policy arguments presented by AI systems.
Authors
Michele Loi
Abstract
An argument can be surprising coming from a particular speaker without being a bad argument. Do AI evaluators keep these judgments apart? Two preregistered descriptive studies and a later Jev supplement collected 2,976 usable evaluations of six fixed texts about US AI policy, Germany's debt brake and Swiss nuclear energy. Each text was presented under several source attributions. The key comparison asks whether the gap between two sources changes when the argument changes. On Sol, for example, a national-security argument received mean ratings of 0.359 under CODEPINK and 0.639 under College Republicans; a civil-rights argument received 0.742 and 0.721. A constant preference for one source cannot explain that pattern. Related interactions appeared across topics and recent model configurations, including those with reasoning enabled, while several comparisons yielded small effects. The later European Jev supplement yielded five interactions below the adopted absolute reference of 0.05; its distinct rubric and interrupted collection limit comparison with the chat systems. Some written evaluations explicitly invoked a mismatch between a source and its attributed position. Taken together, the numerical and verbal evidence supports source-position coherence as a plausible explanation, alongside competing accounts involving credibility, authenticity and interpretation of the task. The paper develops this inference through controlled comparisons, reports conditional post hoc p-values in an appendix, and documents the human decisions and delegated checks behind an AI-conducted study.