Model uses confident answers to stop reading early and save time
The Model Knows When to Stop: Training-Free Early Stopping for Long-Context Reading
Computation and LanguageMachine Learning
Summary
Processing long texts can take lots of computer time if a language model always reads everything fully. The authors create a simple way for a model to decide when it has enough information by checking how confident and stable its answers are, without any extra training. This method, called Answer-Convergence Stopping, saves time by stopping early without losing accuracy. They tested it on tough reading tasks and found it works better than other methods that require extra learning.
What this means in practice
- •For natural language processing teams: Enhance efficiency of processing long textual inputs by dynamically stopping when models produce stable confident answers, lowering computation costs.
- •For customer support automation teams: Reduce response time in automated text understanding systems by stopping the reading once the system is sure of its answer.
Authors
Muath Alyobi, Mohamed Eltahir, Almoayyad Abuljdail, Riyadh Almutawa, Tanveer Hussain, Naeemullah Khan
Abstract
Language models often process long inputs sequentially in chunks, but continuing to read after sufficient evidence has been acquired wastes computation. Existing stopping mechanisms either learn sufficiency from internal activations or train an exit gate, while a simpler alternative asks the model whether it has read enough. We introduce Answer-Convergence Stopping (ACS), a training-free stopping rule that measures rather than asks. After each chunk, it probes the frozen model's current answer state and stops when that state is both confident and stable. The rule requires only output-side generation and token log probabilities, has no trained components, and uses one shared configuration across models and benchmarks. Because a stopping policy can save computation simply by stopping too early, we evaluate the stopping decision itself using evidence position where available. On the full LongBench-v2 with two frontier models, ACS is the only stopping policy that matches or exceeds full-reading accuracy. Furthermore, across 250 S-NIAH questions, the premature stopping rate for ACS across five models from two families ranges from 0% to 12%, compared to 8.4% to 45.6% for the verbalized gate. Taken together, ACS reveals that by properly utilizing the output signals of frozen models, we can achieve favorable behaviors like adaptive stopping without the need for additional training.