Papers for

scientific data engineers

Papers whose findings have a practical use for this group, as judged from the abstract. Open a paper to read what it means in practice.

Schema adaptation improves scientific PDF data extraction accuracy

Adapt Semantics, Not Structure: Few-Instance Schema Calibration for Scientific PDF Extraction

Abstract: A well-designed extraction schema is not necessarily ready for reliable LLM execution. When only limited verified extractions are available, manually tuning hundreds of field definitions through trial and error is costly. We frame this problem as few-instance schema calibration: adapting the operational semantics of an existing schema from a few annotated documents while preserving its structural contract. We introduce CPSE, a contract-preserving semantic extraction framework that jointly calibrates extraction prompts and field-level semantic descriptions from a few gold annotations. CPSE decomposes the schema into an invariant structural contract and mutable field semantics, and further separates identity discovery from record completion using manifest-conditioned resolution. On expert-annotated polymer-science documents, CPSE improves extraction by 9.93 points over an execution-matched baseline, with consistent gains under an independent judge and in a blinded expert audit. These results show that CPSE enables low-resource schema execution while preserving the output structure required downstream.

Mon 28 SeptComputation and Language
The gist
Extracting specific information from scientific PDFs is hard, especially when few examples are available to guide the process. The authors present a method that tweaks the meaning of data fields without changing the overall structure, making the extraction more accurate with just a few examples. Their approach, called CPSE, keeps the output organization stable while improving results on polymer science documents. This helps people get reliable data from complex scientific texts with limited manual effort.
Open → 2609.34841v1

Optimizing scientific data extraction from weak task descriptions

From Weak Task Specifications to Scientific Extraction Agents: Optimizing Task Construction

Abstract: Most methods that optimize LLM prompts and agent workflows assume that task-specific output schemas, extraction instructions, and evaluation criteria are predefined. For scientific extraction agents, however, a short task goal may not fully determine these components, while specifying them manually is costly. We study the upstream problem of constructing the task-specific configuration from a weak specification containing only a short goal and unannotated reference documents. Rather than treating automatic construction as a fixed preprocessing step, our framework constructs a task-specific schema, extraction instructions, and base training rubrics, then keeps schema construction and extraction instructions editable during optimization. Failure-focused updates concentrate textual-gradient feedback on lower-scoring documents, while training-time evaluation criteria adapt to recurring failures. On a heterogeneous-catalysis literature corpus, automatic construction remains improvable, and optimizing both schema construction and extraction instructions performs best across all four judge-rubric settings, with ablations and blinded human evaluation supporting the proposed formulation.

Mon 28 SeptComputation and Language
The gist
Many methods using large language models need detailed instructions and rules for extracting information, but defining these manually is hard and time-consuming. The authors studied how to create these instructions automatically from just a short goal and example texts without annotations. Their approach keeps refining the extraction rules and training checks by focusing on mistakes and learning from them. Testing on chemistry research papers showed the method can improve information extraction by tweaking both the data format and instructions during training.
Open → 2609.34829v1