Legal reward models improve grounded reasoning and abstention in law ai

Building Legal Reward Models for Grounding and Abstention

Computation and LanguageArtificial Intelligence

Summary

Computers used in legal work need to base their answers on real evidence and say "I don't know" when they lack good information. The authors created new ways to train and test models that judge if these AI answers are well grounded in law documents. Their approach improved how the models behave, especially when the evidence is noisy or missing. They also found that training on legal data from one place helps the models do better in different legal systems too.

What this means in practice

  • For legal technology developers: Build AI systems that better evaluate and produce legally grounded answers with reliable abstention when evidence is insufficient.$Commercial implications: This enables production of more trustworthy AI legal assistants and analytics tools for law firms and legal service providers.
  • For ai system trainers: Improve reward model training using context-aware preference data and length-balanced augmentation to enhance retrieval-based generation evaluation.

Authors

Rilton Franzone, Valentin Noël, Puyu Wang, Philip Torr, Fabio J. Fehr

Abstract

Large language models are increasingly used in high-stakes domains such as law, where systems must ground their reasoning in retrieved evidence and abstain when that evidence is insufficient. However, existing reward models are largely optimised for general preferences rather than contextual grounding, limiting their ability to evaluate these behaviours in retrieval-augmented generation (RAG) settings. We introduce a framework for transforming existing legal QA datasets into contextual preference data and use it to construct LegalRewardBench (LRB), a benchmark for evaluating grounded legal generation under noisy and insufficient retrieval conditions. Across general and legal contextual evaluation, we find that contextual DPO improves grounded evaluation, but performance is sensitive to preference-data construction. Length-balanced augmentation substantially improves grounded legal evaluation, with the strongest configuration combining length-balanced legal and general contextual preference data and improving performance by up to $\mathbf{+25.6}$pp over baseline. We further find evidence of cross-jurisdiction transfer: models contextually refined primarily on Victorian criminal-law data improve grounded evaluation on external US legal benchmarks, including a $\mathbf{+16.2}$pp improvement on \textsc{Housing Statute QA}. Together, these results provide a reproducible foundation for constructing and evaluating grounded legal reward models in retrieval-augmented settings.