Benchmark measures large language models on rule based decisions and reasons
RGDT-Bench: Benchmarking LLM Reasoning for Rule-Governed Decisions and Their Justifications
Computation and Language
Summary
People use rules to make decisions in things like contracts and policies, but computers need to do more than just follow rules—they must also explain their reasoning clearly. The authors created a big test called RGDT-Bench that checks if computer programs can correctly apply rules and give good reasons for their decisions. They found that many programs often give incomplete explanations, which can be risky. To fix this, they trained a new model that better detects when explanations are missing or unclear.
What this means in practice
- •For compliance teams: Evaluate and improve automated systems that make and explain decisions in regulatory or policy contexts to reduce safety risks from incomplete justifications.
- •For contract management developers: Test and enhance decision-making AI that must apply complex contract rules and provide checkable reasoning for each decision.
- •For legal technology providers: Develop improved AI tools that justify decisions following legal rules with supervised benchmarks to ensure completeness and consistency of reasoning.$Commercial implications: Enables sale of AI decision support systems to legal firms requiring reliable and justifiable interpretations of rules.
Authors
Jianpeng Zhao, Haihua Xu, Haoyang Zhang, Shuang Qian, Yixiang Tang, Xintao Wang, Kun Sun, Pei Wu, Shuhan Zhong, Pengyang Wang
Abstract
We study reasoning in Rule-Governed Decision Tasks (RGDTs), where models apply external rules to case facts and justify decisions, as required in policy, contract, and compliance settings. Beyond the deductive capability emphasized by standard mathematical and logical reasoning tasks, RGDTs require interpreting rules and their applicability, assessing conditions from evidence, combining judgments under rules and exceptions, and providing checkable justifications. These demands motivate a benchmark assessing both decisions and their stated grounds. We introduce RGDT-Bench, providing 202.1K condition-level supervision slots across four task tracks and eight supported task-probe combinations that vary access to supporting information. Label-blind extraction and deterministic checks produce labels for warrant completeness: source-referenced coverage and consistency of stated decision grounds. The benchmark attributes failures to four process layers: rule use, condition, evidence, and aggregation, and checks the final outcome. Among evaluable correct responses, warrant incompleteness averages 40.2% across six evaluated LLMs and supported task-probe combinations. Such warrant incompleteness poses potential safety risks and remains difficult to detect: the best of seventeen existing evaluators reaches only 57.69% (random: 50%) task-averaged area under the receiver operating characteristic curve (AUROC). To address this difficulty, we train a simple reward model with warrant supervision. It achieves 69.24% task-averaged AUROC among correct answers, exceeding the matched outcome-supervised baseline by 10.37 pp (percentage points) and the best existing evaluator by 11.55 pp. Beyond completeness assessment, the model outperforms both outcome-supervised baselines across nearly all response-selection comparisons, supporting RGDT-Bench's warrant supervision for RGDT reasoning.