Forge improves large language model answers while cutting token costs

FORGE: Form-Optimal Routing of Grounded Evidence for Frozen LLM Agents

Machine LearningComputation and Language

Summary

Using big AI models as fixed services limits how they can be improved because you can't change their internal settings. The authors found that picking what information to show and how much time the model spends thinking usually isn’t done well with fixed settings. They created FORGE, a system that decides these choices for each question to get better answers while using fewer words. The system learns how to make these decisions without changing the original AI model, and it works well across many tasks and model types.

What this means in practice

  • For software engineers: Implement efficient question-answering systems that adapt reasoning effort and evidence form per query to reduce costs and improve accuracy with fixed AI model APIs.
  • For ai platform operators: Optimize usage of frozen language model endpoints by dynamically routing input formats and inference budgets, lowering expenses while maintaining performance.
  • For customer support teams: Enhance automated response quality and reduce cost by tailoring evidence provision and processing depth per customer query using frozen AI backends.$Commercial implications: Enables commercial AI-powered support tools to deliver better and cheaper answers by customizing model interaction strategies.

Authors

Xi Xiao, Yunbei Zhang, Chen Liu, Lin Zhao, Jialin Chen, Tianchen Zhao, Xiang Xu, Youngeun Kim, Tianyang Wang, Min Xu

Abstract

In agentic AI systems, frozen foundation models are increasingly deployed as closed-weight API endpoints, making downstream adaptation possible only through the inputs and inference procedures surrounding the model. As a result, for each input query, two coupled decisions largely determine both answer quality and token cost: what evidence to provide and how much reasoning budget to allocate. Fixed defaults along these axes are often suboptimal, misallocating support form or reasoning depth on roughly 80% of queries in our analysis. To address this challenge, we propose FORGE, a unified framework for adapting frozen models through per-query routing over a joint action space that spans both support form and thinking depth. Under an entropy-regularized, cost-aware utility objective, we derive a closed-form Boltzmann routing target and instantiate the policy as a lightweight 269K-parameter factorized router. The routing policy is trained around the frozen host, without any weight access, through a three-stage pipeline: offline arm enumeration, supervised Kullback-Leibler (KL) distillation from the Boltzmann target, and Group Relative Policy Optimization (GRPO) refinement with host feedback. Across 5 knowledge-intensive benchmarks and 8 frozen backbones ranging from 7B to 671B parameters, FORGE improves accuracy at 42-45% lower token cost on both main hosts, transfers zero-shot across hosts at lower token cost, and composes with intrinsic thinking budgets where available.