CLIN: an Objective Framework for Evaluating Creativity in Short Persian Literary Text
2026-08-31 • Computation and Language
Computation and Language
AI summaryⓘ
The authors studied how well large language models (LLMs) can judge creativity in short Persian texts, which is hard because creativity has many parts and depends on human feelings. They found LLMs agree better with humans on clear traits like Originality and Fluency, but not on feelings like Emotion. Different ways of asking the LLMs affected results, but more complicated prompting didn’t help much. So, the authors created a simpler method, called CLIN, that uses clear measures to judge things like novelty and word diversity, matching human judgment as well as complex LLM methods but much faster and cheaper.
Large Language ModelsCreativity EvaluationTTCTPersian LanguageOriginalityFluencyElaborationPrompt EngineeringLexical DiversityNovelty Metrics
Authors
Mohammad Reza Modarres, Armin Tourajmehr, Yadollah Yaghoobzadeh, Mohammad Taher Pilehvar
Abstract
Evaluating creativity in large language model (LLM) outputs remains challenging because creativity is multidimensional and human-centered. We examine how reliably LLMs evaluate short literary text in Persian, a low-resource language, across multiple evaluation strategies and prompt formulations. We find that LLM-human agreement varies substantially across dimensions: alignment is stronger for structured TTCT-derived properties such as Originality, Fluency, and Elaboration, but considerably weaker for more subjective dimensions, particularly Emotion and Attractiveness. Judgments are also sensitive to prompt formulation, while few-shot prompting, ensembling, and multi-agent debate provide no consistent improvement. Motivated by this dimension-dependent behavior, we investigate whether structured creativity dimensions can instead be approximated using simple, interpretable proxy metrics. We introduce CLIN, which evaluates three TTCT-derived dimensions separately using topic-aware novelty for Originality, contextual lexical clustering for Fluency, and lexical diversity for Elaboration. These proxies achieve human alignment comparable to or better than the strongest zero-shot LLM judge in our setting while requiring substantially lower evaluation cost.