Source-preserving alignment improves locating evidence text in scientific pdfs
Source-preserving alignment for robust evidence localization in scientific PDFS
Artificial Intelligence
Summary
Finding exact evidence in scientific papers is hard because the text displayed can differ from how it’s stored, with line breaks and special characters causing confusion. The authors present a new way to match pieces of evidence text back to their original spots in PDFs, keeping track of where the text came from while allowing some flexibility in matching. This method works much better than simple text search or older techniques, especially on chemistry papers. It helps verify scientific claims by showing exactly where in a paper the supporting text is found.
What this means in practice
- •For legal document reviewers: Identify and precisely highlight text evidence in complex formatted PDFs to verify claims and references quickly.
- •For pharmaceutical research teams: Accurately locate supporting evidence within chemical research PDFs to speed up literature validation and data extraction.
Authors
Zihao Liu, Wei Yang, Zixiao Dong, Chenshu Li, Longzhang Liu, Tao Tan, Hong Xie
Abstract
Scientific information-extraction systems often return a claim with an evidence string, which users must locate in the original PDF. This is challenging because the extracted evidence and PDF text layer are different representations: line wrapping, Unicode variants, superscripts, citation markers, and fragmented items alter text sequences and geometry. We present a source-preserving alignment framework: normalize text for robust matching while preserving provenance for accurate localization. It aligns evidence with normalized page text, maps matches back to source-character spans, and renders only their geometry. When exact alignment fails, line-break-aware token alignment recovers supported spans while excluding unmatched noise. Experiments on 1,020 chemistry papers show that the framework achieves a 92.6\% quote-level automatic localization rate, compared with 43.6\% for text search and 19.1\% for a precomputed bounding-box baseline. Component ablation confirms distinct contributions from normalization and approximate token alignment, while human verification assesses the visual correctness of returned highlights. Overall, these results demonstrate that reliable evidence verification requires robust matching and precise localization within a shared source-preserving alignment representation.