EviStreams lets medical teams control AI data extraction for reviews

EviStreams: Human-in-the-Loop AI Data Extraction for Systematic Reviews in Medicine

Computation and Language

Summary

Systematic reviews in medicine require experts to carefully pull data from many studies, which is time-consuming and must follow strict rules. The authors present EviStreams, a web platform that helps review teams guide AI in extracting data while keeping human oversight and follow protocols to ensure accuracy. Experts use a form builder to define data fields, review AI-suggested values alongside evidence, and resolve any disagreements anonymously for a reliable final dataset. They found that how fields are defined impacts extraction quality more than which AI model is used.

What this means in practice

  • For medical review teams: Enable teams to control AI-assisted data extraction from medical studies while preserving strict review protocols and auditability.
  • For legal document reviewers: Help teams extract structured data from complex legal texts in a protocol-driven, auditable way using dual review and adjudication.

Authors

Sai Karthik Kosuri, Ankita Shashikant Bhosale, Michael Glick, Alonso Carrasco-Labra, Chris Callison-Burch

Abstract

Systematic reviews underpin clinical guidelines, yet their data-extraction step is a major expert-labor bottleneck bound by a protocolized workflow: two reviewers extract each study independently, an adjudicator resolves disagreements, and the team keeps an auditable record of how every value was produced. Large language models can assist with extraction, but that assistance must fit established review protocols and preserve reproducibility. We present EviStreams, a live, open-source, no-code web platform that puts review teams in control of AI-assisted extraction at three key stages: program design (a structured decomposition approved before any code runs), field specification (typed field definitions calibrated from a pilot), and extracted predictions (reviewer-blinded dual review with adjudication). Working through a form builder, a domain expert defines typed fields rather than prompts, runs extraction over uploaded PDFs, inspects every value alongside the supporting passage it came from, and resolves a reviewer-blinded dual review into an auditable consensus export. An evaluation across four clinical corpora and three frontier model families, released with the system, shows that extraction quality is shaped far more by the field specification than by the choice of model. EviStreams is live at https://evistreams.com/demo and released under Apache-2.0.