Dual-axis optimization method improves learning for large language model agents
Dual-Axis Policy Optimization for LLM Agents: Bayesian Feedback Attribution and Trajectory Mass Normalization
Artificial Intelligence
Summary
Training large language model agents to perform tasks based on feedback is complicated because it involves two steps: figuring out which parts of their actions led to feedback inside a single attempt, and combining results across many attempts. The authors introduce a new way called BATON that handles these two steps separately but together. One part uses Bayesian methods to better assign credit to actions inside a single attempt, while the other balances how many complete attempts are considered equally during training. Tests show that using both techniques together helps agents learn better on different tasks and models.
What this means in practice
- •For llm developers: Train language model agents more effectively by improving how learning feedback is assigned and combined across training attempts.
- •For ai platform engineers: Enhance reinforcement learning pipelines with batch normalization techniques that treat full agent behaviors equally during optimization.
Authors
Yingxuan Zhuang, Binhe Yu, Jingxiao Yang, Ruopei Sun, Ziting Li, Cheng Tan, Xuhong Zhang, Jianwei Yin, Jintao Chen
Abstract
Reinforcement learning for LLM agents involves two distinct optimization di- mensions: how environment feedback is exploited within a trajectory, and how complete trajectories are aggregated across a batch. We formulate these dimen- sions as Intra-Trajectory Feedback Attribution and Inter-Trajectory Objec- tive Aggregation, and introduce BATON (Bayesian Attribution and Trajectory Objective Normalization), a dual-axis policy optimization framework. BATON instantiates the first axis with Bayesian Feedback Attribution, which constructs a feedback-conditioned posterior over sampled actions, and the second with Trajec- tory Mass Normalization (TMN), which assigns equal optimization mass to com- plete trajectories. Experiments with GRPO and GiGPO on ALFWorld, WebShop, and SearchQA show that both axes provide independent gains and that their combi- nation consistently achieves the strongest overall performance across model scales.