Update-aware method improves multi-domain language agent training

USA: Update-aware SAM for Cross-domain On-Policy Disitllation of Language Agents

Computation and Language

Summary

Training AI language agents to reason in many areas is hard because mixing data from different topics can cause conflicts and poor learning. The authors found that simply combining knowledge from separate fields often hurts performance due to conflicting updates on the model’s parameters. They propose a method called USA that measures how much each model parameter changes and protects those most affected, allowing better merging of knowledge from multiple domains. Tested in areas like math, science, and coding, this approach improves learning and reverses previous negative effects.

What this means in practice

  • For language model developers: Improve multi-domain reasoning models by independently training on each domain then merging updates without negative interference across topics.
  • For software engineering teams: Enhance code generation tools by combining specialized knowledge from math, science, and programming domains with less retraining.

Authors

Qiyong Zhong, Mao Zheng, Mingyang Song, Huwei Ji, Houcheng Jiang, Jiajie Su, Li Zhang, Gengsheng Li, Junfeng Fang

Abstract

On-policy distillation instils multi-turn agentic reasoning through dense token-level supervision on the student's own trajectories, but a single domain saturates early, so further supervision has to be drawn from other domains. Multi-domain data mixing is the most direct way of incorporating them, at the cost of conflicts between their data distributions and of retraining the entire model whenever one domain is revised. Model merging avoids both by distilling every domain independently and fusing the resulting task vectors afterwards. We find instead that the benefit polarizes across domain pairs: on those exhibiting negative transfer, every merging operator we evaluate falls below the single-domain reference. We attribute this to cross-domain update coupling, where a substantial fraction of coordinates is updated comparably by both domains and a merge can therefore displace them by as much as their own updates. To overcome this limitation, we propose USA, which converts per-parameter update magnitudes measured during a brief warm-up into per-coordinate perturbation radii, reducing curvature precisely on the coordinates that carry most of the merging displacement. Experiments across mathematics, science and code at two student scales show USA strongest in all six transfer directions, ahead of the single-domain reference by more than four points on average, and reverse the negative transfer of the conflicting pairs.