Papers for

digital archive managers

Papers whose findings have a practical use for this group, as judged from the abstract. Open a paper to read what it means in practice.

Multi-vector visual document search sped up with smaller query models

ColNanoVDR: Document-Free Query Distillation for Multi-Vector Visual Document Retrieval via Optimal Transport

Abstract: Multi-vector retrievers built on vision-language models lead visual document retrieval (VDR), but they run a multi-billion-parameter query encoder on every search. Distilling this encoder into a small student that queries the teacher's existing index would remove the bottleneck. The standard recipe, however, matches the teacher's MaxSim scores and so requires encoding and caching every training page, which can reach terabytes of page tokens. NanoVDR avoids pages entirely by training on the teacher's query embeddings alone, but only for single-vector retrievers. We present ColNanoVDR, to our knowledge the first framework to bring this document-free distillation to multi-vector VDR. Its objective, OTW (Optimal Transport with Learned Weights), aligns the student's query tokens with the teacher's by entropic optimal transport, with a learned weight for each student token, and needs no correspondence between the two tokenizations. We prove that the resulting alignment cost bounds the MaxSim score difference on every page. Distilled from five state-of-the-art teachers, the 149M text-only students retain about 95% of their teachers' NDCG@5 on ViDoRe v1-v3 while encoding queries up to 26x faster. Under identical training, OTW matches score distillation while encoding no page and reading 12.6x less cached teacher data.

Mon 28 SeptInformation RetrievalComputation and LanguageComputer Vision and Pattern Recognition
The gist
Searching large collections of visual documents is slow because the computer runs a huge model every time. The authors developed a way to teach a smaller, faster model to ask the big model for help without needing to look at all the documents. They use a math technique called optimal transport to match parts of the smaller model’s questions to the big model’s answers. This method nearly keeps the search accuracy while making query processing up to 26 times faster.
Open → 2609.34899v1

Historical persian manuscript dataset improves word spotting accuracy

HPMD: A Historical Persian Manuscript Dataset for Word Spotting with Line-Level Annotation

Abstract: Large collections of historical Persian manuscripts have been digitized, but searching them is still slow and mostly manual. Historians usually want to find where a specific name, date, event, or topic appears, which is a word spotting problem. Progress on this task is limited by two things. First, there is almost no public dataset of historical Persian handwriting; the only notable resource, OpenITI MAKHZAN, contains a relatively small Persian portion. Second, word spotting models usually need word-level bounding boxes, which are very expensive to annotate. In this paper we introduce a new dataset of 223 pages, 3,678 lines, 37,631 words, and 130,630 characters, collected from diverse historical Persian books of poetry and prose and annotated at the region, line, and text level. We also propose a baseline that is trained only with line-level annotations but returns word-level locations. A fine-tuned line detector finds text lines, and a fine-tuned CRNN recognizer trained with CTC produces a frame-by-character posterior matrix for each line. Instead of decoding the most probable character at each frame, the query is scored directly against this matrix, so visually similar characters in Persian such as be and pe no longer cause hard failures. The frame alignment also gives the horizontal position of the word inside the line. On the test set, the fine-tuned line detector reaches an F1 of 0.892, and posterior-based search raises the word spotting F1 from 0.487 to 0.558 compared with exact matching on the decoded text, with the decision threshold selected on a held-out validation set. A PHOC attribute-embedding baseline that additionally receives oracle word boundaries at test time reaches an F1 of 0.449, below the proposed method. We also report a distributional analysis of the dataset, a taxonomy of retrieval errors, and a per-conditionbreakdown of performance.

Mon 28 SeptComputer Vision and Pattern Recognition
The gist
Searching old Persian handwritten books for specific words is hard because there are no good public datasets and marking words manually is expensive. The authors created a new large dataset of Persian manuscripts with detailed line and text labels. They also developed a method that finds lines and spots words without needing word-level marking during training, improving search accuracy. This approach handles visually similar Persian letters better, making it easier to locate words in the text.
Open → 2609.34490v1

Visual document retrieval improves by adapting queries with residual feedback

Test-Time Adaptation with Query-Dependent Residuals for Visual Document Retrieval

Abstract: Visual document retrieval (VDR) systems depend on page embeddings computed before deployment, which makes adaptation difficult when encoder parameters or corpus re-encoding are unavailable. Rerankers provide useful relevance signals, but conventional reranking applies them only to selected queries and candidate pages. We introduce Q-REACT, a query-side test-time adaptation method that converts limited reranker feedback into reusable retrieval improvements. Q-REACT learns a shared low-rank transformation that produces query-dependent residuals, combines adapted query scores with document-level context, and distills reranker preferences with a student distribution normalized over the complete task-specific page index. This design lets unscored pages compete through cached embeddings while keeping the encoders and page index fixed. Across eight ViDoRe V3 tasks and five open-weight and proprietary backbones, Q-REACT improves average retrieval over evaluated baselines at sparse and full-coverage budgets, transfers to held-out queries and tasks, and adds little inference overhead. The results show that finite reranker feedback can be amortized across a query collection without retraining or rebuilding the retriever.

Wed 23 SeptInformation Retrieval
The gist
Visual document retrieval systems find relevant pages in large document collections using precomputed page data, but adapting them after deployment is hard. The authors present Q-REACT, a method that uses feedback from a secondary ranking process to adjust queries without changing the underlying document data. This approach allows all documents to compete fairly during retrieval and works well across multiple tasks and retrieval systems. It improves performance with minimal extra computing time and does not require changing or re-encoding the documents.
Open → 2609.27688v1

Query adaptive indexing improves retrieval in expert archives

ORDER: Task-Conditioned Routing for Retrieval-Augmented Generation

Abstract: Retrieval-Augmented Generation (RAG) pipelines typically rely on a fixed indexing and retrieval configuration determined at preprocessing time. This one-size-fits-all design is ill-suited to domain-expert settings, where heterogeneous queries require different chunking granularities, metadata constraints, and source-selection strategies. As a result, configurations that are effective for one family of queries often perform poorly for others. In this paper, we introduce ORDER (Optimal Routing for Dynamic Evidence Retrieval), a query-conditioned RAG framework that jointly adapts indexing and retrieval to the incoming query. Our approach first discovers semantic clusters over a given set of questions associated to a corpus and learns, for each cluster, a chunking strategy together with a suited metadata filtering and reranking configuration. At inference time, queries are routed to the appropriate pre-built index through nearest-centroid assignment. To further improve retrieval, we propose a supervised query router (QRe) that predicts which collections are most likely to contain relevant evidence, coupled with a Uniform Multi-source Sampler (UMS) that allocates the retrieval budget evenly across the selected sources. We evaluate our framework on large-scale, heterogeneous historical archives and show that conditioning both indexing and retrieval on the query consistently outperforms both naive baselines and strong state-of-the-art RAG systems in complex expert-domain environments.

Tue 15 SeptArtificial Intelligence
The gist
Many systems that find and generate answers from large collections use the same setup for all questions, even though different questions need different ways of organizing and searching the data. The authors propose ORDER, a method that changes how the data is divided and searched depending on the type of question asked. First, it groups similar questions and learns the best way to split and filter the underlying information for each group. Then, when a new question comes in, it routes it to the right way of searching based on its similarity to the groups. This approach helps get better answers from complex, historical archives with diverse information.
Open → 2609.17012v1