Language models learn to improve their task strategies at run time
Harness Learning Enables Generalizable Test-Time Adaptation
Computation and LanguageMachine Learning
Summary
Modern AI language models work with instructions called harnesses that tell them how to organize their actions and tools for different tasks. This paper presents a way for the model to learn how to improve these instructions by trying them out and getting feedback, without changing its internal model itself. This lets the model adapt to new tasks it hasn’t seen before by revising how it solves problems during use. The authors show that this approach helps the model get better at reasoning and answering questions that require multiple steps.
What this means in practice
- •For ai developers: Enable language model applications to adapt task-solving procedures dynamically for improved reasoning on new, unseen tasks at run time.
- •For automated customer support teams: Improve chatbot response quality by fine-tuning dialogue control strategies during conversations without retraining the language model parameters.
Authors
Alvin Zhang, Xuecheng Liu, Zixuan Wang, Fahim Tajwar, Daman Arora, Ruslan Salakhutdinov, Daniel Khashabi, Yuda Song, Andrea Zanette
Abstract
A language-model agent is jointly defined by its model and its harness, the executable program that organizes model calls, tool use, and information flow. Because different tasks call for different ways of organizing these operations, the harness needs to be adapted using feedback from the task at hand. We introduce harness learning, which trains a proposer model to revise a solver's harness using execution feedback. We formulate this process as meta-learning over executable programs, with harness revisions playing the role of weight updates in gradient-based adaptation. We train the proposer with reinforcement learning, using the task performance of revised harnesses as the reward. At test time, the proposer uses feedback from successive executions on a new task to refine the harness, without performing any parameter-space update. Experiments on reasoning and multi-hop question answering show that harness learning improves revision quality and that the ability to adapt at test time transfers to unseen tasks. Policies trained on individual revisions can continue improving harnesses over multiple rounds, while the benefits of training on revision sequences vary across settings. These findings suggest a path towards continually learning agents that turn accumulated experience into generalizable improvements.