Optimizing scientific data extraction from weak task descriptions

From Weak Task Specifications to Scientific Extraction Agents: Optimizing Task Construction

Computation and Language

Summary

Many methods using large language models need detailed instructions and rules for extracting information, but defining these manually is hard and time-consuming. The authors studied how to create these instructions automatically from just a short goal and example texts without annotations. Their approach keeps refining the extraction rules and training checks by focusing on mistakes and learning from them. Testing on chemistry research papers showed the method can improve information extraction by tweaking both the data format and instructions during training.

What this means in practice

  • For scientific data engineers: Automatically generate and optimize information extraction templates from minimal task descriptions to speed up scientific literature processing.
  • For knowledge management teams: Improve document processing workflows by dynamically refining extraction instructions based on error feedback during training.

Tested on one dataset.

Authors

Zixiao Dong, Wei Yang, Zihao Liu, Chenshu Li, Longzhang Liu, Tao Tan, Hong Xie

Abstract

Most methods that optimize LLM prompts and agent workflows assume that task-specific output schemas, extraction instructions, and evaluation criteria are predefined. For scientific extraction agents, however, a short task goal may not fully determine these components, while specifying them manually is costly. We study the upstream problem of constructing the task-specific configuration from a weak specification containing only a short goal and unannotated reference documents. Rather than treating automatic construction as a fixed preprocessing step, our framework constructs a task-specific schema, extraction instructions, and base training rubrics, then keeps schema construction and extraction instructions editable during optimization. Failure-focused updates concentrate textual-gradient feedback on lower-scoring documents, while training-time evaluation criteria adapt to recurring failures. On a heterogeneous-catalysis literature corpus, automatic construction remains improvable, and optimizing both schema construction and extraction instructions performs best across all four judge-rubric settings, with ablations and blinded human evaluation supporting the proposed formulation.