QuanReview system improves accuracy of annotation comparisons

QuanReview: Offline, Auditable Reconciliation of Human and LLM Span Annotations

Computation and Language

Summary

Annotations in documents, like numbers with units or event types, can be tricky and costly to create accurately. The authors present QuanReview, a tool that checks and fixes differences between human and AI-made annotations by comparing them closely and letting humans resolve unclear conflicts. This helps keep annotations trustworthy and saves time by automatically agreeing on simple cases. It was tested with a humanitarian dataset and showed it reduced human work while improving data quality.

What this means in practice

  • For data annotation teams: Automate and streamline reviewing differences between human and AI annotations to improve data quality in annotation projects.
  • For humanitarian data analysts: Ensure trustworthy extraction of sensitive quantitative and event data from humanitarian reports by reconciling human and model annotations.

Authors

Matteo Musacchio, Juan Cruz Giner Pulero, Isabel Castañeda, Naomi Couriel, Yelena Mejova, Mariano G. Beiró, Kyriaki Kalimeri

Abstract

Structured span annotations, such as quantities with their units, uncertainty modifiers, and event classes, are expensive to create and hard to keep trustworthy once language models enter the loop. We present QuanReview, an open-source system for auditing and correcting such annotation layers. QuanReview aligns two annotation streams over the same documents at character level, resolves unambiguous cases by an explicit and logged policy, and routes candidate conflicts to a browser-based adjudication interface where reviewers accept either side, build field-level hybrids, or flag items for re-annotation. A campaign manager assigns documents to multiple annotators with configurable redundancy, computes agreement at document and span level, auto-merges unanimous documents, and exports the corrected layer in the original file format, so that it can replace the original annotation files directly. Applied to a 4,457-record humanitarian benchmark and an LLM extraction stream, the system fully auto-merged 8% of documents, applied automatic policy decisions to a further 1,513 records, and concentrated human attention on 3,131 candidate conflicts, a mean of 5.4 per reviewed document.