LogicTrack audits reasoning steps of language models with logic solvers
LogicTrack: Auditing Reasoning Trajectories of Large Language Models with Formal Logic Solvers
Artificial IntelligenceLogic in Computer ScienceSymbolic Computation
Summary
Large language models can explain their answers step-by-step, but those steps might have mistakes even if the final answer is right. The paper introduces LogicTrack, a tool that checks each reasoning step by turning it into formal logic and verifying it automatically. It also helps fix mistakes during reasoning by backtracking when errors are found. This method improves the trustworthiness of language model answers, especially in important areas where errors matter.
What this means in practice
- •For ai system developers: Build language models that provide logically verified reasoning steps to increase trust in AI outputs for critical decisions.
- •For software testing teams: Automatically audit the intermediate reasoning logic of AI systems to detect and correct reasoning flaws before deployment.
- •For legal technology providers: Use step-wise logic verification in AI tools that assist with legal reasoning to ensure sound and auditable outputs.$Commercial implications: Enables development of AI-assisted legal decision tools with validated step-by-step reasoning trusted by clients and courts.
Authors
Jingyu Hu, Shu Yang, Weiru Liu, Di Wang
Abstract
Chain-of-Thought (CoT) reasoning has been shown to improve the performance of large language models (LLMs), yet existing optimization methods largely rely on outcome-based feedback, leaving the logical validity of intermediate reasoning steps largely unverified. To address the gap whereby LLMs arrive at correct final answers through logically flawed intermediate reasoning chains, we propose LogicTrack, a neuro-symbolic framework that audits reasoning trajectories by auto-formalizing each reasoning step into symbolic representations and verifying it with automated theorem provers. LogicTrack introduces Solver-Based Backtracking Reward (SBR), a step-wise scoring mechanism that quantifies logical soundness and guides backtracking tree search at inference time. We further extend LogicTrack to construct supervised fine-tuning (SFT) data with backtracking traces from its trajectories, enabling fine-tuned models to internalize step-wise auditing as an intrinsic capability. Extensive experiments across 8 reasoning benchmarks and 7 LLMs demonstrate that LogicTrack effectively improves both the verifiability of reasoning chains and final answer pass rate, thereby enhancing overall CoT quality and trustworthiness in high-stakes domains.