Papers for

legal technology providers

Papers whose findings have a practical use for this group, as judged from the abstract. Open a paper to read what it means in practice.

Benchmark measures large language models on rule based decisions and reasons

RGDT-Bench: Benchmarking LLM Reasoning for Rule-Governed Decisions and Their Justifications

Abstract: We study reasoning in Rule-Governed Decision Tasks (RGDTs), where models apply external rules to case facts and justify decisions, as required in policy, contract, and compliance settings. Beyond the deductive capability emphasized by standard mathematical and logical reasoning tasks, RGDTs require interpreting rules and their applicability, assessing conditions from evidence, combining judgments under rules and exceptions, and providing checkable justifications. These demands motivate a benchmark assessing both decisions and their stated grounds. We introduce RGDT-Bench, providing 202.1K condition-level supervision slots across four task tracks and eight supported task-probe combinations that vary access to supporting information. Label-blind extraction and deterministic checks produce labels for warrant completeness: source-referenced coverage and consistency of stated decision grounds. The benchmark attributes failures to four process layers: rule use, condition, evidence, and aggregation, and checks the final outcome. Among evaluable correct responses, warrant incompleteness averages 40.2% across six evaluated LLMs and supported task-probe combinations. Such warrant incompleteness poses potential safety risks and remains difficult to detect: the best of seventeen existing evaluators reaches only 57.69% (random: 50%) task-averaged area under the receiver operating characteristic curve (AUROC). To address this difficulty, we train a simple reward model with warrant supervision. It achieves 69.24% task-averaged AUROC among correct answers, exceeding the matched outcome-supervised baseline by 10.37 pp (percentage points) and the best existing evaluator by 11.55 pp. Beyond completeness assessment, the model outperforms both outcome-supervised baselines across nearly all response-selection comparisons, supporting RGDT-Bench's warrant supervision for RGDT reasoning.

Mon 28 SeptComputation and Language
The gist
People use rules to make decisions in things like contracts and policies, but computers need to do more than just follow rules—they must also explain their reasoning clearly. The authors created a big test called RGDT-Bench that checks if computer programs can correctly apply rules and give good reasons for their decisions. They found that many programs often give incomplete explanations, which can be risky. To fix this, they trained a new model that better detects when explanations are missing or unclear.
Open → 2609.34455v1

LogicTrack audits reasoning steps of language models with logic solvers

LogicTrack: Auditing Reasoning Trajectories of Large Language Models with Formal Logic Solvers

Abstract: Chain-of-Thought (CoT) reasoning has been shown to improve the performance of large language models (LLMs), yet existing optimization methods largely rely on outcome-based feedback, leaving the logical validity of intermediate reasoning steps largely unverified. To address the gap whereby LLMs arrive at correct final answers through logically flawed intermediate reasoning chains, we propose LogicTrack, a neuro-symbolic framework that audits reasoning trajectories by auto-formalizing each reasoning step into symbolic representations and verifying it with automated theorem provers. LogicTrack introduces Solver-Based Backtracking Reward (SBR), a step-wise scoring mechanism that quantifies logical soundness and guides backtracking tree search at inference time. We further extend LogicTrack to construct supervised fine-tuning (SFT) data with backtracking traces from its trajectories, enabling fine-tuned models to internalize step-wise auditing as an intrinsic capability. Extensive experiments across 8 reasoning benchmarks and 7 LLMs demonstrate that LogicTrack effectively improves both the verifiability of reasoning chains and final answer pass rate, thereby enhancing overall CoT quality and trustworthiness in high-stakes domains.

Fri 18 SeptArtificial IntelligenceLogic in Computer ScienceSymbolic Computation
The gist
Large language models can explain their answers step-by-step, but those steps might have mistakes even if the final answer is right. The paper introduces LogicTrack, a tool that checks each reasoning step by turning it into formal logic and verifying it automatically. It also helps fix mistakes during reasoning by backtracking when errors are found. This method improves the trustworthiness of language model answers, especially in important areas where errors matter.
Open → 2609.21492v1

Question answering system excels with version and scope awareness in legal texts

Version- and Scope-Aware Question Answering over Normative Documents: A Deployed System and an End-to-End Evaluation at Production Scale

Abstract: Correctly answering a question grounded in normative documents often depends on information outside any single passage: whether the retrieved document is the version currently in force; whether it applies to the jurisdiction, subject (such as an institution or applicant), and date at issue; and whether each normative claim can be traced to its supporting source text. Hosted retrieval services have substantially lowered the engineering cost of building an initial system over such corpora, making "upload the documents and ask" a common default. We evaluate this default on approximately 73,000 candidate normative documents supplied to a production deployment. The evaluation uses a stratified sample of 200 questions from our published benchmark, with a gold source document for every question; the released sampling rule reads no system outputs or scores. We compare the hosted service with a governed system that resolves version and scope through explicit rules before generation. The governed system scored 97.7 overall, while the hosted service scored 88.1, a gap of 9.6 points computed from unrounded means. The question set, the answer text evaluated for both systems, the scores, and the scripts used to reproduce the reported benchmark statistics are public. The governed configuration has operated as a commercial product since January 2026 and serves 1,126 registered users; named customer organizations include Zhipu AI and Lecheng Health. By mid-April 2026, it had reached roughly 100,000 calls per workday.

Wed 16 SeptArtificial Intelligence
The gist
Answering questions about legal or official documents is tricky because it matters which version of the document is current, where it applies, and who it affects. The authors compared two systems: one that simply searches documents without extra checks, and one that carefully follows rules about document versions and who they apply to before answering. Their system that understands versions and scope answered questions more accurately. This system is already in commercial use, helping over a thousand users with tens of thousands of queries daily.
Open → 2609.18769v1