Schema adaptation improves scientific PDF data extraction accuracy
Adapt Semantics, Not Structure: Few-Instance Schema Calibration for Scientific PDF Extraction
Computation and Language
Summary
Extracting specific information from scientific PDFs is hard, especially when few examples are available to guide the process. The authors present a method that tweaks the meaning of data fields without changing the overall structure, making the extraction more accurate with just a few examples. Their approach, called CPSE, keeps the output organization stable while improving results on polymer science documents. This helps people get reliable data from complex scientific texts with limited manual effort.
What this means in practice
- •For scientific data engineers: Improve data extraction from specialized scientific PDFs with minimal annotated samples by refining field meanings instead of redesigning schemas.
- •For document processing teams: Increase reliability of automated information extraction in workflows that handle complex, structured PDF documents using limited verified extraction examples.
Authors
Zixiao Dong, Wei Yang, Zihao Liu, Chenshu Li, Longzhang Liu, Tao Tan, Hong Xie
Abstract
A well-designed extraction schema is not necessarily ready for reliable LLM execution. When only limited verified extractions are available, manually tuning hundreds of field definitions through trial and error is costly. We frame this problem as few-instance schema calibration: adapting the operational semantics of an existing schema from a few annotated documents while preserving its structural contract. We introduce CPSE, a contract-preserving semantic extraction framework that jointly calibrates extraction prompts and field-level semantic descriptions from a few gold annotations. CPSE decomposes the schema into an invariant structural contract and mutable field semantics, and further separates identity discovery from record completion using manifest-conditioned resolution. On expert-annotated polymer-science documents, CPSE improves extraction by 9.93 points over an execution-matched baseline, with consistent gains under an independent judge and in a blinded expert audit. These results show that CPSE enables low-resource schema execution while preserving the output structure required downstream.