STONIC: A Layered Measurement Contract for LLM Value Profiling

2026-08-24Computation and Language

Computation and Language
AI summary

The authors studied how large language models (LLMs) express their preferences through ratings, choices, and generated text to see if these all show the same stable preference. They tested many models and situations and found that while some patterns hold across tasks, models often prefer their own earlier answers and are influenced by factors like option position. Preferences seen in ratings are strongest when compared to choices made under conflict but weaker in spontaneous text generation. The authors also checked how well different evaluation methods agree with human judgments and found mixed results. Overall, the models display some consistent behavior, but the evidence suggests preferences are not exactly identical across different ways of measuring them.

large language modelspreference elicitationpairwise choicerating scalessemantic auditbehavioral continuitymodel calibrationcounterbalanced conflictfixed model configurationsprompt engineering
Authors
Andrei Chetvergov, Stepan Ukolov, Timofei Sivoraksha, Alexander Evseev, Danil Sazanakov, Mikhail Solovev, Sergey Bolovtsov
Abstract
LLM value studies often merge questionnaire ratings, pairwise choices, and values inferred from generated text into one profile. That merge assumes that the three observations describe the same stable preference. STONIC tests this assumption on 5,144 situations from four banks and 35 fixed model configurations. It compares responses rated in isolation, choices made under counterbalanced conflict, spontaneous answers, and later choices between a model's own answer and authored alternatives. 10 of 17 configurations with usable behavioral data preserve the endorsement-choice relation across banks. Every one of the 17 eligible configurations prefers its own earlier answer (median effect 0.790), although option position changes the choice rate in every eligible configuration. Profile shape transfers most strongly from ratings to conflict choices and weakens for spontaneous text. Three-way annotation of 200 L3 responses provides a task-local check of the semantic audit: FULCRA agrees most closely with the human majority, while DeBERTa retains useful rank information after calibration. Hidden states encode the completed decision more clearly than the prompt alone. Thus the models show reproducible behavioral continuity, but the evidence does not support one scorer-independent value identity across interfaces.