Harness-zero transfers specialized agent skills into model weights
Harness-Zero: Harness Distillation via Agent-as-Harness
Artificial IntelligenceComputation and LanguageNeural and Evolutionary Computing
Summary
Different tasks and environments need different helpers for AI agents, which can slow things down or limit performance if you stick to one helper. The authors study a way to teach an AI model the best behaviors created by task-specific helpers, so those skills stay even when the special helper is removed. They do this by having a helper agent fix the AI's output before it acts, creating examples for the AI to learn from. Their method, called Harness-Zero, improves AI performance significantly and retains useful behaviors from specialized helpers without needing those helpers during use.
What this means in practice
- •For ai developers: Improve AI agents' task success by embedding specialized helper skills into models for use without extra systems.
- •For automation engineers: Deploy AI agents that maintain complex task behaviors without needing diverse external control modules.
Authors
Haoran Ye, Yuxing Lu, Haonan Dong, Zhaochen Su, Guojie Song
Abstract
Agent harnesses, the external systems that mediate model-environment interaction, can substantially improve agent performance, but their gains remain tied to the harness at deployment. Because the best harness varies across domains, instances, and models, a general-purpose agent must either settle for a suboptimal shared harness or route among an ever-growing set of specialized ones. We therefore study agent harness distillation: using a domain- or instance-optimized harness as training-time guidance and transferring the behaviors it induces into model weights, so that its gains survive under a single fixed target harness. The challenge is that the two harnesses differ in action space and available information, so guidance from the optimized harness cannot serve directly as supervision for the target one. We introduce Harness-Zero, which enables harness distillation through agent-as-harness. Guided by the optimized harness, a harnessing agent corrects student responses before execution in the target harness's action space, turning harness guidance into training demonstrations. Fine-tuning on the resulting trajectories internalizes harness-induced behavior into the model, so the specialized harness can be removed at deployment. Our experiments spanning knowledge work, tool use, and science domains show that: (1) For frontier LLMs using the same evolved harness, agent-as-harness outperforms code-as-harness. (2) With the specialized harness removed at deployment, Harness-Zero improves the base model's macro-average task success from 23.3% to 44.3%, even exceeding the 41.7% it reaches with that harness still attached. (3) Harness-Zero recovers harness-induced behaviors absent from the base model, with 82.3% average recovery across 28 patterns in the three domains.