Question answering system excels with version and scope awareness in legal texts

Version- and Scope-Aware Question Answering over Normative Documents: A Deployed System and an End-to-End Evaluation at Production Scale

Artificial Intelligence

Summary

Answering questions about legal or official documents is tricky because it matters which version of the document is current, where it applies, and who it affects. The authors compared two systems: one that simply searches documents without extra checks, and one that carefully follows rules about document versions and who they apply to before answering. Their system that understands versions and scope answered questions more accurately. This system is already in commercial use, helping over a thousand users with tens of thousands of queries daily.

What this means in practice

  • For legal technology providers: Provide accurate answers from multiple legal documents by verifying current versions and relevant jurisdiction before responding to user queries.$Commercial implications: Enables sale of advanced legal AI tools to firms needing reliable legal compliance answers, improving accuracy over generic search solutions.
  • For healthcare compliance teams: Assist compliance officers by delivering precise regulatory answers tied to correct document versions and applicable subjects in health regulations.

Authors

Liuyin Wang, Shuaipeng Jin, Jiwei Shi, Jensen Hsu

Abstract

Correctly answering a question grounded in normative documents often depends on information outside any single passage: whether the retrieved document is the version currently in force; whether it applies to the jurisdiction, subject (such as an institution or applicant), and date at issue; and whether each normative claim can be traced to its supporting source text. Hosted retrieval services have substantially lowered the engineering cost of building an initial system over such corpora, making "upload the documents and ask" a common default. We evaluate this default on approximately 73,000 candidate normative documents supplied to a production deployment. The evaluation uses a stratified sample of 200 questions from our published benchmark, with a gold source document for every question; the released sampling rule reads no system outputs or scores. We compare the hosted service with a governed system that resolves version and scope through explicit rules before generation. The governed system scored 97.7 overall, while the hosted service scored 88.1, a gap of 9.6 points computed from unrounded means. The question set, the answer text evaluated for both systems, the scores, and the scripts used to reproduce the reported benchmark statistics are public. The governed configuration has operated as a commercial product since January 2026 and serves 1,126 registered users; named customer organizations include Zhipu AI and Lecheng Health. By mid-April 2026, it had reached roughly 100,000 calls per workday.