Papers for

document management teams

Papers whose findings have a practical use for this group, as judged from the abstract. Open a paper to read what it means in practice.

Page aware retrieval improves French PDF question answering accuracy

Page-Aware Retrieval-Augmented Generation for EvalLLM 2026: A Five-Variant Study on French PDFs

Abstract: We study retrieval-augmented generation (RAG) for questions about French PDF documents when both the answer and its supporting document pages are evaluated. Five system variants add dense retrieval, rank fusion, reranking, and query decomposition to a BM25 baseline. On 595 challenge questions, the complete system scores 0.4450 MRR@10 and 0.4013 Recall@10, compared with 0.3430 and 0.2994 for BM25. Dense retrieval alone and a simple lexical--dense fusion both underperform BM25. Reranking improves the hybrid system, whereas adding query decomposition produces the largest further gain, with higher latency and more detected output artifacts. The complete system slightly exceeds the reported anonymous overall mean on two answer metrics but falls below it on most page-retrieval metrics. These results identify accurate page selection, rather than semantic retrieval in isolation, as the main opportunity for improvement in this setting.

Mon 28 SeptArtificial Intelligence
The gist
Finding the right page in French PDF documents is key to answering questions accurately. The authors tested five ways to combine search methods, starting with a basic keyword search called BM25. They found that mixing methods and breaking down queries helped pick better pages and answers, though it made the system slower and caused some errors. This suggests improving page selection is more important than just understanding meanings better.
Open → 2609.34776v1

RidgeRank speeds up visual document reranking with score fusion

RidgeRank: Efficient Visual Document Reranking via Score Fusion and a Shallow Linear Readout

Abstract: Multimodal language models rerank visual document retrieval results accurately, but scoring every candidate page at full cost makes them slow. Some methods that compress these rerankers need relevance labels to regain accuracy, and they rank by the reranker score alone. RidgeRank measures how much relevance signal the reranker score lacks and recovers it from the retriever score through a closed-form fusion rule. Maximizing a correlation objective gives the optimal fusion weight, along with the exact condition under which the reranker score by itself cannot reach that optimum. The reranker is further corrected by a single vector applied to an intermediate hidden state, obtained through one centered ridge regression onto the same model's full-depth scores on uncompressed pages. On 12 datasets drawn from ViDoRe 2 and ViDoRe 3, evaluated with two retrievers and two language model backbones, RidgeRank brings NDCG@5 to within 1.2 pp of a full cross encoder with speedups of up to 48 times, advancing the accuracy and latency Pareto frontier for visual document reranking.

Mon 28 SeptInformation Retrieval
The gist
Sorting visual documents to find the most relevant ones can be slow when using powerful language models. The authors propose RidgeRank, a way to quickly combine scores from a fast retriever and a slower, more accurate reranker. By cleverly mixing these scores, RidgeRank keeps most of the accuracy but runs much faster. Tests on multiple datasets show it nearly matches full reranking accuracy with several times speedup.
Open → 2609.34192v1

Vision language models improve long document question answering pipelines

An Empirical Study of VLM Pipelines for Long-Document QA

Abstract: Vision-Language Models (VLMs) are increasingly used for long-document processing, where the inputs combine text with charts, tables, figures, and complex layouts. Deploying them means choosing how to feed the document to the model, which retriever to use when only a subset of pages is sent, and whether to run the model agentically or as a static pipeline. We study these choices on two long-document QA benchmarks with both frontier API and open-weight VLMs. First, on MMLongBench-Doc our six-tool agent with page, table, figure, and search calls pays off only once the answering VLM is large enough: with Qwen3.5-4B and 9B it trails static page input, with Qwen3.5-27B it draws level, and with Sonnet 4.5 it leads. On LongDocURL it is level with or ahead of static input at every reader. Its lead over the strongest static pipeline is clearest with the frontier reader on MMLongBench-Doc and narrows to within noise on LongDocURL. Second, retrieval modality matters more than the specific retriever: the strongest image retriever leads the strongest text pipeline, and on the text side a single off-the-shelf cross-encoder rerank essentially matches a much heavier multi-stage LLM pipeline. Top-k image retrieval is also the most token-efficient input at every reader we paired it with, at roughly a seventh to a quarter of the tokens of sending every page. Third, cutting across all three choices, three of our strongest pipelines succeed on different questions, and an oracle that picks the best pipeline per question gains roughly thirteen points over the best single pipeline, though evidence-type routing recovers almost none of it.

Thu 24 SeptComputation and LanguageArtificial IntelligenceComputer Vision and Pattern Recognition
The gist
Long documents often contain text plus charts, tables, and images, which makes it hard for AI to answer questions about them accurately. The authors compare different ways to feed these documents into vision-language AI models and test multiple strategies on two benchmarks. They find that models that handle images well and use selective retrieval do better and use fewer words as input. Also, combining different pipeline methods could answer more questions, although it’s challenging to pick the best method automatically.
Open → 2609.29933v1

Text aware single step method improves text image resolution

TOLA: Text-aware One-Step Latent Adaptation for Diffusion-based Text Image Super-Resolution

Abstract: Text image super-resolution (TSR) aims to recover visually faithful and readable text under unknown degradations. Existing diffusion-based methods typically rely on multi-step prediction of either the high-resolution image or its text prior, resulting in prohibitive computational cost and inference latency. More critically, an erroneous text prior may be repeatedly injected into the denoising process, causing image and text predictions to reinforce each other and progressively amplify an early recognition error into a sharp yet semantically incorrect character. To address these limitations, we propose TOLA, a Text-aware One-step Latent Adaptation framework without iterative image-text diffusion. TOLA consists of two key modules. First, a confidence-weighted text conditioning module constructs the semantic condition only once and suppresses unreliable OCR predictions before they contaminate image reconstruction. Second, a lightweight latent residual correction module explicitly estimates and corrects the structured residual errors to recover missing or distorted stroke details. Extensive experiments demonstrate our state-of-the-art performance across all evaluation metrics on both CTR-TSR-Test ($\times 4$) and RealCE-200 benchmarks. It is worth noting that our TOLA consistently surpasses existing diffusion-based TSR methods by at least 2.72 dB in PSNR on CTR-TSR-Test.

Thu 24 SeptComputer Vision and Pattern RecognitionArtificial Intelligence
The gist
Text image super-resolution tries to make blurry or damaged text in images clear and readable again. Previous methods took many steps and were slow, sometimes making errors worse by repeating wrong guesses. The authors propose a new method called TOLA that works in just one step and avoids repeating mistakes by carefully judging which text guesses to trust. Their method also fixes small details in the text strokes to create clearer letters. Tests show TOLA works better and faster than previous approaches.
Open → 2609.29240v1

AI detectors struggle to spot fake document images accurately

Beyond Natural Images: Rethinking AI-Generated Image Detection in Documents

Abstract: AI-generated image detection has attracted increasing attention, but existing evaluations mainly focus on natural images, leaving AI-generated document images largely underexplored. This omission is concerning because documents often appear in sensitive real-world scenarios, such as invoices, expense reports, certificates, and medical records. In this paper, we first construct a controlled diagnostic benchmark, AIGDoc-Pilot, and reveal that existing detectors suffer substantial performance degradation on AI-generated document images, with the mean AUC dropping by more than 7%. Based on this, we further reveal two document-specific properties behind this gap: generation artifacts exhibit strong spatial inconsistency across local regions, and text density significantly affects real-synthetic separability, where text-dense regions offer stronger discriminative evidence. Motivated by these findings, we construct AIGDoc, a larger document-centric dataset containing diverse real-world documents and AI-generated counterparts produced by multiple advanced generation and editing models. Extensive experiments on AIGDoc demonstrate that existing detectors still struggle to reliably identify AI-generated documents, while document-based training partially narrows the gap. Together, these results offer valuable insights for developing dependable and generalizable detectors in document-centric scenarios. The code and datasets will be made publicly available upon acceptance of the paper.

Sun 13 SeptComputer Vision and Pattern Recognition
The gist
Detecting images made by AI usually works well on photos, but not on document images like invoices or certificates. The authors found current detection tools perform worse on AI-generated documents because of unique features like uneven artifact patterns and text density. They created a new dataset of real and AI-made documents and showed that training detectors specifically on documents can help improve detection. This work is important for stopping fake documents from causing problems in sensitive areas.
Open → 2609.14352v1

Visual document search improves storage with generative embeddings

Generative Late-Interaction Embeddings For Visual Document Retrieval

Abstract: Late-interaction retrieval is the state-of-the-art for visual document search, but it pays for its accuracy in storage. Existing compression methods retain a subset or local average of the N~1,000 vectors per page. Under aggressive storage budgets, however, these methods degrade sharply, and alternatives require retraining the encoder. Investigating this degradation across three encoders, we found two consistent properties: the vectors lie exactly on the unit sphere and concentrate near a manifold of intrinsic dimension five to six. This geometry yields two insights. First, standard k-means centroids fall inside the sphere, causing systematic underestimation of MaxSim scores. Normalizing them to the surface is a free correction worth up to +0.093 nDCG@5 over raw centroids. Second, because the page manifold has few degrees of freedom, the full set of vectors can be regenerated from only a few. To this end, we introduce Generative Late-Interaction Embeddings (GLIE): k << N vectors per page learned from the normalized centroids to serve as both a lightweight index and a basis for regenerating the page's full embedding set. At query time, search runs exclusively on these k vectors, and a decoder expands only the top candidates back to all N vectors for exact rescoring. At four vectors per page on ViDoRe v1, GLIE retains nearly 80% of the uncompressed system's nDCG@5, against 70% for the best prior post-hoc method. These results use a 415K-parameter network fitted in under three GPU-minutes on just a thousand training pages. At a matched training budget, fine-tuning the encoder does not reach even the training-free stage of GLIE, and the full system beats it at every budget. These patterns hold across a second encoder and ViDoRe v2. By reconstructing evidence on demand rather than sampling it, GLIE opens a new axis for storage-efficient retrieval, with the decoder as its main design surface.

Thu 10 SeptInformation Retrieval
The gist
Searching one thousand visual data vectors per page uses a lot of space and slows tools down. The authors discovered that these vectors cluster near a shape with only five to six dimensions and lie on a fixed-radius sphere. They created a method called GLIE that stores just a few summary vectors and can recreate the full detailed set when needed. This approach keeps most of the search accuracy but drastically reduces storage needs and speeds up retrieval.
Open → 2609.11808v1

Cassette improves legal case search with faster and accurate retrieval

Cassette: Case-to-Case Structural Distillation for Efficient Legal Case Retrieval

Abstract: Legal case retrieval (LCR) is an essential tool for not only assisting legal practitioners to efficiently retrieve precedents but also enabling ordinary individuals to find valuable legal case information without relying on expensive professional legal services. Our previous work CaseLink demonstrated the effectiveness of using case to case graph structures to improve retrieval accuracy. However, its high computational cost during inference on large-scale legal databases limits its practical use in real-world settings. The main inefficiency comes from constructing test time graphs and computing pairwise term frequency similarities of cases. This process has O(n^2) complexity for n legal cases, making the runtime prohibitive as the number of candidates grows. For example, the retrieval time for one query on a database (COLIEE2022) with 1,563 candidate cases is more than 500 milliseconds, while the runtime would increase drastically to more than 3,500 seconds for a database (LeCaRDv2) with 55,192 candidate cases. To further enhance the retrieval performance while achieving a significant speed-up, in this extension paper, Cassette framework is proposed with a distillation strategy involving ranking objective and eigen-matching objective for an effective transfer of knowledge from a powerful and well-trained heavy teacher retriever to a lightweight and efficient hybrid student dual encoder. Specifically, the student query encoder is implemented as a multilayer perceptron model designed for fast online processing, whereas the student candidate encoder adopts a GNN architecture, suitable for an offline manner within the case database. Extensive experiments are conducted on three benchmark datasets, and the results verify the effectiveness of the ranking distillation while achieving high efficiency. The code has been released on https://github.com/yanran-tang/Cassette/.

Tue 8 SeptInformation Retrieval
The gist
Finding relevant legal cases quickly can be very slow when the database is large. The authors built on previous methods that linked cases to each other but were too slow for big collections. They designed a new system called Cassette that learns from a complex model and transfers this knowledge to a faster, simpler one. This makes searching through thousands of cases much faster while still being accurate. They tested their method on several datasets and confirmed it works well.
Open → 2609.08185v1