Papers for

document ai developers

Papers whose findings have a practical use for this group, as judged from the abstract. Open a paper to read what it means in practice.

Compact model with program harness matches large models for document math

When Harness Beats Scale, and When Reading Beats Both

Abstract: We describe our system for DocSem, the document-grounded quantitative reasoning shared task at DocInsights 2026, and analyze why it succeeded on labeled data and failed on the test set. The pipeline pairs hybrid block retrieval with Program-of-Thoughts (PoT) generation executed in a sandboxed interpreter, self-consistency sampling, and entity enrichment from chunk-level knowledge graphs. On our held-out split, application architecture moved the metrics far more than model scale did: PoT added 0.282 joint accuracy to a compact 7B model but at most 0.005 to a 72B model, and a 27B model with the full harness matched the 72B (0.884 vs.\ 0.873) at roughly 2.7$\times$ fewer parameters and a quarter of the CO$_2$. We read this through a distinction between world knowledge, which scales steeply with parameters, and language knowledge, which scales gently, and show that structured-output training makes a compact model harness-ready rather than merely small. On the raster, watermarked test PDFs the same system collapsed to 13.58\% joint (rank 149 of 163); a controlled re-rendering of the validation set reproduces the OCR half of the collapse while bounding what the simulation misses. Auditing the physical nature of evaluation inputs precedes architecture, and the leaderboard's bimodality is consistent with reading quality, not reasoning, having separated the field.

Mon 28 SeptComputation and LanguageArtificial IntelligenceInformation Retrieval
The gist
Solving math problems based on documents usually relies on very large AI models, which are costly and slow. The authors show that using a clever setup—called a program-of-thoughts harness—helps smaller models perform just as well as much bigger ones on labeled data. They also found that problems with reading documents from low-quality PDFs caused many systems to fail in tests, showing that how well a system reads input matters more than its reasoning ability. This means improving reading quality can be more important than making models bigger.
Open → 2609.34366v1

Page images improve document QA accuracy but increase latency and cost

Text, Pixels, or Both? Evaluating Input Representations for Multimodal Document QA

Abstract: Every document QA system begins with a choice that is rarely studied on its own: whether to feed the model page images, extracted text, or both. We isolate this choice, holding the prompt, judge, and scoring pipeline fixed, across four commercial model endpoints, two corpora, and two context regimes (gold evidence pages and the full document). On documents that fit the image budget, page images lead on accuracy at every document length on both corpora, but this advantage carries a growing latency and cost premium: text latency stays roughly flat as documents lengthen while image latency rises steadily. Text and images also fail on different questions, with exactly one representation correct on 19--25% of items across the reported cells, so neither subsumes the other. Exploiting this complementarity, a lightweight TF-IDF router that reads only the question text gains 2.6 points over always-text while cutting median latency 30% relative to always-vision, on a document-disjoint held-out split.

Fri 18 SeptArtificial Intelligence
The gist
When computers answer questions about documents, they can use either the words extracted from the pages, images of the pages themselves, or both. This paper shows that using page images generally helps computers answer more accurately but takes more time and computing power as documents get longer. Interestingly, some questions are answered correctly only by looking at images or only by looking at text, so using both can be better. The authors also built a simple system that chooses between text or image input based on the question, improving speed and accuracy.
Open → 2609.22628v1