Solver guided rewards improve logical reasoning in language models

Rewarding Novel Deductions: Solver-guided Process Rewards for Logical Reasoning

Computation and Language

Summary

Large language models struggle with logical reasoning problems that require many careful thinking steps. The authors created a method called SPRING that uses a computer solver to check each step the model takes, rewarding new valid thoughts and discouraging repeated or contradictory ones. This makes the model better at solving puzzles by thinking more clearly and avoiding mistakes. They tested SPRING on several logic puzzle datasets and with different language models, finding big improvements in puzzle-solving accuracy.

What this means in practice

  • For software developers: Improve AI assistants that solve logic puzzles or reason systematically by integrating solver-based rewards during training to enhance step-by-step deduction.
  • For game designers: Build more challenging and accurate AI characters in logic-based puzzle games by using models trained with solver-guided reasoning step rewards.

Authors

Muhammad Asif Ali, Wenqing Wang, Huan Wang, Mohammad Raza

Abstract

Logical reasoning remains a major challenge for large language models (LLMs), particularly on structured problems that require precise constraint tracking, consistency preservation, and multi-step deduction. This challenge is especially acute for small-scale LLMs, which are more prone to producing inconsistent, redundant, or brittle reasoning trajectories. Existing approaches for improving logical reasoning largely optimize for final-answer correctness, providing only weak supervision over the intermediate reasoning process. In this work, we propose SPRING: (Solver-guided Process Rewards for Novel LogIcal ReasoNing Step Generation). SPRING uses SMT solver as a training-time verifier of intermediate reasoning steps to provide process-level supervision. It introduces the notion of a novel reasoning step, namely, a step that is logically valid, consistent with the evolving reasoning state, and not already implied by previously accepted non-contradictory deductions. Based on this solver-based assessment, it designs process rewards that encourage novel inferential progress while penalizing contradictory and uninformative reasoning steps. Evaluation across three logical reasoning benchmarks, ZebraLogic, AR-LSAT, and Knights and Knaves, and four LLMs shows that SPRING consistently outperforms base LLMs, outcome-only reward baselines, and Logic-LM. On ZebraLogic, SPRING improves puzzle accuracy by up to 49.71 and 15.43 points over the base LLM and strongest outcome-only baseline, respectively. On AR-LSAT, it improves overall accuracy by up to 64.93 and 12.14 points, respectively. On Knights and Knaves, SPRING achieves up to 93.14 puzzle accuracy and 96.05 person accuracy.