LEX-EC: A Lexical Evidence-Channel Audit Framework for Zero-Shot LLM Personality Classification in Black-Box Settings
2026-07-27 • Computation and Language
Computation and LanguageArtificial Intelligence
AI summaryⓘ
The authors created a tool called LEX-EC to better understand how large language models decide personality traits from text. This tool checks how reliable the models are when certain words or topics are removed, helping to see if the models rely on meaningful trait signals or just common patterns. They found that different text types, like essays or social media posts, vary in how much personality info they show. The study also shows some word types are more important for identifying traits. Overall, the tool helps explain what influences the model's personality guesses without looking inside the model.
large language modelspersonality labelingmodel interpretabilitylexical ablationprevalence diagnosticsagreement diagnosticstrait associationsprompt sensitivityfunction wordsaffective terms
Authors
Brittany Harbison, Ashok K. Goel
Abstract
Large language models may easily assign personality labels from text, but model interpretability remains an open problem. To address this gap, we introduce LEX-EC, a reusable black-box audit framework combining prevalence and agreement diagnostics with controlled lexical ablation to distinguish marginal-distribution effects from trait-associated signal recoverable under restricted evidence. Using this framework, we illustrate how various text genres may exhibit sharply different profiles: free-form essay text contains the broadest, but still weak, signal; in graduate student introductions, an observable Extraversion association weakened after masking; and single Facebook statuses yield little stable evidence even in a trait-balanced sample, indicating a possible lower bound of content or length. Masking topical and demographic content weakened some associations while leaving others detectable from function words, affective terms, and cognitive-style vocabulary. Linguistic prompting shifted model self-explanations but did not eliminate topical content. LEX-EC jointly evaluates classification prevalence, item-level association, chance-corrected agreement, persistence under lexical restriction, and prompt sensitivity in model-generated explanations. Across datasets, models, and prompts, LEX-EC characterizes how trait associations may vary with available lexical evidence, introducing a novel application of lexical methods to black-box interpretability in personality labeling.