Hidden state trajectories improve detection and repair of reasoning errors in diffusion language models
LOCKR: A Hidden-State Trajectory-Guided Planner for Detecting and Repairing Stable-but-Wrong Lock-In in Diffusion Language Models
Computation and Language
Summary
Sometimes, AI models that write answers get stuck on wrong ideas early and don’t change them later, even though they keep thinking. The authors studied how to spot and fix these mistakes by looking inside the model’s thinking process, not just the final answer or how confident it seems. They created a method called LOCKR that checks the model’s hidden thinking path and decides when to try other options to fix errors. This method improved the model’s accuracy on math reasoning tasks by a few percentage points and fixed many wrong answers.
What this means in practice
- •For language model developers: Improve language model output accuracy by detecting and selectively repairing mistaken early lock-ins during text generation.
- •For automated math solving platforms: Increase reliability of mathematical answers by using hidden-state guided planning to identify and correct stable but wrong intermediate solutions.
Authors
Guoshenghui Zhao, Tan Yu, Weijie Zhao
Abstract
Diffusion language models generate text through iterative denoising, exposing intermediate trajectories before final answers are produced. We identify a recurring reasoning failure, stable-but-wrong lock-in, where an answer stabilizes early around an incorrect value while substantial denoising remains. Surface-level decoding signals such as confidence, entropy, margin, and answer stability are insufficient to reliably distinguish correct from erroneous lock-in. We formulate selective reasoning repair as a lightweight test-time planning problem and propose LOCKR, a hidden-state trajectory-guided planner that decides when to allocate additional computation, expands a structured set of targeted repair branches, and selects the most promising continuation using trajectory-aware verification. Across two diffusion language models and three mathematical reasoning benchmarks, hidden-state trajectories consistently outperform surface signals and single hidden snapshots for both wrong-lock-in detection and repair selection. On natural evaluation distributions, LOCKR yields absolute accuracy gains of 2.21--5.37 percentage points across all five evaluated settings, with repair rates ranging from 22% to 41%. These results establish hidden diffusion trajectories as actionable signals for selective test-time reasoning repair.