Papers for

legal support teams

Papers whose findings have a practical use for this group, as judged from the abstract. Open a paper to read what it means in practice.

Page aware retrieval improves French PDF question answering accuracy

Page-Aware Retrieval-Augmented Generation for EvalLLM 2026: A Five-Variant Study on French PDFs

Abstract: We study retrieval-augmented generation (RAG) for questions about French PDF documents when both the answer and its supporting document pages are evaluated. Five system variants add dense retrieval, rank fusion, reranking, and query decomposition to a BM25 baseline. On 595 challenge questions, the complete system scores 0.4450 MRR@10 and 0.4013 Recall@10, compared with 0.3430 and 0.2994 for BM25. Dense retrieval alone and a simple lexical--dense fusion both underperform BM25. Reranking improves the hybrid system, whereas adding query decomposition produces the largest further gain, with higher latency and more detected output artifacts. The complete system slightly exceeds the reported anonymous overall mean on two answer metrics but falls below it on most page-retrieval metrics. These results identify accurate page selection, rather than semantic retrieval in isolation, as the main opportunity for improvement in this setting.

Mon 28 SeptArtificial Intelligence
The gist
Finding the right page in French PDF documents is key to answering questions accurately. The authors tested five ways to combine search methods, starting with a basic keyword search called BM25. They found that mixing methods and breaking down queries helped pick better pages and answers, though it made the system slower and caused some errors. This suggests improving page selection is more important than just understanding meanings better.
Open → 2609.34776v1

Large language models enhanced with clear stepwise decision making

Integrating the Analytic Hierarchy Process with Large Language Models for Transparent Multi-Criteria Decision-Making

Abstract: LLMs are increasingly employed in a wide range of decision-making tasks. However, the opacity of their internal reasoning makes it difficult to validate or interpret their outputs, and the need for interpretability becomes especially critical in high-stakes settings. This study examines the decision-making capabilities of LLMs through the Analytic Hierarchy Process (AHP), a classical and widely used multicriteria decision-making framework. We construct a new annotated benchmark based on AHP and propose the first end-to-end approach that enables LLMs to perform the complete AHP workflow. Experiments in real-world decision problems in the legal and higher-education ranking domains show that our method significantly improves alignment with expert judgments.

Tue 15 SeptArtificial Intelligence
The gist
Decisions made by large language models can be hard to understand or trust, especially when these models explain their choices in ways that are not clear. The authors combined these models with the Analytic Hierarchy Process, which is a step-by-step method for making choices by comparing different factors carefully. They created a new dataset to teach and test the models on this method and found their approach better matched expert opinions in complex areas like legal cases and university rankings. This work helps make AI decisions more transparent and reliable.
Open → 2609.16779v1