Generative AI shows promise but falls short automating data extraction
Can Generative AI Automate Data Extraction for Meta-Analysis? A Case Study on Intercropping Research
Artificial Intelligence
Summary
Meta-analysis combines information from many studies to find big-picture answers, but pulling data from papers is slow and tricky. The authors tested how well large language models (LLMs) can automatically dig out data on crop mixtures from scientific articles. They found that while one method did best, none were perfectly accurate, and the automatic results only roughly matched human conclusions. This study shows that LLMs could help but aren’t ready to replace people for detailed data extraction yet.
What this means in practice
- •For agricultural research teams: Assist teams with preliminary data extraction to accelerate meta-analyses on intercropping research despite imperfect accuracy.
- •For environmental data analysts: Provide a method to quickly identify general trends in mixed crop studies for environmental impact assessments.
Authors
Zehao Lu, Xingguo Xiong, Wopke van der Werf, Thijs L. van der Plas, Ioannis N. Athanasiadis
Abstract
Meta-analysis is the synthesis of information from multiple sources to arrive at an overarching conclusion. There is a large need for meta-analysis in agricultural research to synthesize what is known and analyze overarching patterns. Extracting data from published literature is, however, labor-intensive, time-consuming, and tedious, and is impeded by a lack of standardization in research design, units of measurement, and terminology. These challenges are particularly evident in the domain of crop species mixtures, also called intercropping. With the growing capabilities of LLMs, many recent attempts have focused on building systems and tools to automate data collection, yet rigorous assessment against human-labeled ground truth is often missing. In this research, we evaluate three LLM-based approaches---direct zero-shot prompting, a staged workflow, and a multi-agent system---with six open-weight models to extract data from the intercropping literature. The results are evaluated against the manually curated ground truth and through a downstream statistical analysis. Overall, direct zero-shot prompting is the strongest and most consistent approach, achieving the highest mean similarity-adjusted F1 of 0.577, although none of the approaches is close to fully accurate. In the downstream analysis, most model--approach combinations recover the direction of the relationship between the predictor and outcome variables, but do not estimate its magnitude accurately.