A Dual-Dimensional LLM Framework for Automated Item Incidental Content Similarity Analysis in Large-Scale Assessments
2026-08-25 • Artificial Intelligence
Artificial Intelligence
AI summaryⓘ
The authors address the issue of accidental repeated content in large tests, which can happen when test questions share similar wording or context but shouldn't. They create a new method using advanced language models to measure how alike questions are, focusing on both structure and meaning. Their approach matches better with known testing problems and organizes questions more sensibly than older methods. When tested in adaptive exams, their method helps make better question choices, improving accuracy without much loss in efficiency. Overall, the authors show that using language model-based similarity can improve how large test question banks are managed and used.
Large Language ModelsItem SimilarityAutomatic Item GenerationConstruct-Irrelevant VariancePsychometricsComputerized Adaptive TestingLocal DependenceTest AssemblyBLEU ScoreSemantic Relatedness
Authors
Jing Huang, Jihong Zhang, Hua-Hua Chang
Abstract
The rapid expansion of large-scale assessments and the growing adoption of automatic item generation have intensified concerns about incidental content redundancy, where construct-irrelevant elements such as wording or contextual framing become unintentionally repetitive across items. Traditional similarity metrics like BLEU or cosine similarity, often fail to capture the nuanced structural and semantic layers that drive perceived redundancy simultaneously. This study proposes a dual-dimensional framework for Automated Item Similarity Analysis (AISA) powered by Large Language Models (LLMs), operationalizing similarity through Structured Decomposition and Semantic Relatedness. Psychometric validation indicates that LLM-derived metrics align more closely with indicators of construct-irrelevant local dependence and yield more coherent item parameter groupings than traditional text-based measures. The framework is further evaluated through its application in Computerized Adaptive Testing (CAT). Simulations reveal that incorporating LLM-based similarity constraints into item selection improves estimation stability and reduces bias with minimal efficiency trade-offs, outperforming constraints based on conventional metrics. These findings highlight the potential of LLM-powered AISA to support scalable bank curation, content-aware test assembly, and experience-sensitive adaptive testing across diverse assessment contexts.