Agentic meta reasoning improves long task control in large ai agents
Thinking Before Thinking: Scaling Agentic Inference Through Meta-Reasoning
Artificial Intelligence
Summary
Large AI agents solving complicated problems need to decide how to work step-by-step, like picking what to build on next or when to stop. The authors introduce agentic meta-reasoning, which adds a smart controller that manages these decisions instead of just following fixed steps. This controller summarizes progress, explores options, evaluates their value, and chooses the best next move without rereading everything done before. Their tests show this approach helps AI perform better on complex programming, reasoning, and proof tasks by reusing earlier work and making better choices.
What this means in practice
- •For software development teams: Improve AI-assisted coding tools by giving the AI better control over multi-step program writing tasks for more accurate software reconstruction.
- •For automated reasoning engineers: Enhance long-horizon reasoning systems by adopting meta-reasoning controllers to better allocate limited compute during stepwise problem solving.
Authors
Paras Dahal, Anton Bakhtin, Taco Cohen, Zhengxing Chen, Carole-Jean Wu, Rob Fergus, Scott Yih, Gabriel Synnaeve, Ruslan Salakhutdinov, Sanjeev Arora, Jason Weston, Anirudh Goyal
Abstract
As agents take on longer and more complex problems, controlling the execution becomes a task in its own right. Each step in the run brings new control choices, like which partial work to build on, whether to start fresh, or when to stop. We introduce agentic meta-reasoning, an inference-time harness that makes these choices an explicit and structured reasoning process. Workers carry out the task-level computation, while a controller consolidates what the run has established, explores next options, assesses what each option is worth under the remaining budget, and dispatches the chosen work with context drawn from persistent memory. Between decisions the controller carries only a compact account of the run rather than replaying its full history. Our baselines span production coding agents and research harnesses, together with a Direct Control Agent using the same workers and compute budget allowance. On ProgramBench, which tests long-horizon agentic capability through program reconstruction, meta-reasoning achieves 71.5% with GPT-5.5 against 58.0% for Codex; with Opus 4.8 it achieves 67.2% against 65.5% for Claude Code. On the other benchmarks, spanning abstract reasoning, multi-domain long-horizon reasoning, and proof generation, it gains between 3.6 and 4.2 points over direct control, averaged across three frontier models. It keeps improving over the tested budget ranges where direct control plateaus, though its overhead can hurt at small budgets. Artifact-graph analysis reveals more reuse of earlier work, higher coverage of correct solutions in most settings, and nonuniform gains in final selection. These results indicate that spending computation on structured control becomes more important as agents scale to longer runs.