AI coders debate and agree to improve qualitative coding accuracy

How AI Coders Discuss, Disagree, and Reach Consensus: Challenges and Opportunities for LLM-Based Qualitative Coding

Human-Computer InteractionArtificial Intelligence

Summary

Figuring out how AI can help people label text data is tricky because it depends on many things. The authors studied how multiple AI agents code, argue, and come to agreements when labeling text from different sources. They found that longer instructions, how similar the data is, and how much the AI agents disagree affect how well the AI does. Surprisingly, harder debates between AI agents sometimes meant the final answers were more accurate. The authors also noticed AI behaves a bit like humans in discussions but doesn't adapt well to changing contexts.

What this means in practice

  • For social science research teams: Improve automated coding workflows by integrating AI agents that discuss and resolve disagreements to enhance coding accuracy on qualitative datasets.
  • For business data analysts: Use AI agents to independently code and debate survey or customer feedback data, helping reach more accurate thematic insights efficiently.

Authors

Jeongyeon Kim, John Mitchell

Abstract

The utility of AI in multi-coder qualitative coding has been widely discussed, yet little empirical evidence exists to delineate the contexts in which it performs reliably. We address this gap by quantifying the effectiveness of multi-agent LLM coding across varied qualitative datasets, revealing key contextual and structural factors that mediate coding outcomes. We developed a literature-informed baseline pipeline that enables AI agents to independently code, debate, and reconcile disagreements. Results revealed that coding accuracy depends on factors such as codebook length, qualitative data similarity, and agent disagreement. Notably, intense and unresolved debates between agents led to higher accuracy. Our analysis showed that while LLMs emulate many human discussion behaviors, they lack adaptive responsiveness to context. From these findings, we offer design recommendations for building automated coding systems. Our open-source AI discussion dataset and methodological framework lay the groundwork for advancing the design of AI-mediated automated thematic analysis.