Agent calibration preserves capabilities across models and markets

From Migration to Calibration: Preserving Agent Capabilities across Models, Jurisdictions, and Scale

Artificial IntelligenceMultiagent SystemsSoftware Engineering

Summary

When software agents are moved to new situations like different countries or new tasks, they often need adjusting to keep working well. The authors explain that simply making sure they fit the new system isn't enough; their abilities must also remain effective. They propose a method to calibrate agents on three levels: keeping useful information, adapting the way they interact with their environment, and meeting user-specific output needs. This approach aims to make sure the agents don't lose their skills and even improve when adapting to new settings. The proposal includes ideas for how to test and evaluate these adjustments, but it has not yet been tried in practice.

What this means in practice

  • For e-commerce platform engineers: Maintain consistent agent behavior when deploying AI models across different country-specific markets with varied rules and customer expectations.
  • For enterprise software developers: Adapt AI agents to new internal tools or datasets while preserving validated capabilities and meeting team-specific output formats.

A position paper. It proposes an approach and reports no results.

Authors

Yaxiao Liu, Pengbo Liu, Yiwen Liu, Yihua Guan, Jiaxing Song

Abstract

Agents need calibration when deployment conditions change: replacing a driving model, including a foundation-to-post-trained transition; crossing jurisdictions; or scaling across heterogeneous markets and sources. Interface compatibility alone does not establish capability retention or target-contract satisfaction. We formulate agent calibration as constrained behavioral adaptation across three interacting layers: information preservation, harness adaptation, and user acceptance; the layers apply to every scenario, not one-to-one to the three. The basic objective is non-degradation on prespecified capability measures while satisfying target requirements; aggregate improvement is stronger. Information calibration preserves independently validated source content still applicable to the target task. Harness calibration aligns observable artifacts at semantic checkpoints and repairs them through iteration, tool substitution, or local replanning within explicit budgets. User calibration enforces recipient-specific output contracts: templates, schemas, and section-level preferences. A global e-commerce example shows how shared standards coexist with site- and market-specific adapters and validation. We distinguish trainable policies from frozen-backbone configuration or controller optimization, and evidence verification from relative judgment and DPO/GRPO optimization. Recent harness-transfer and judge-validity studies motivate target-native execution records, separate audits of task validity and near-tie ranking, and matched target-native optimization controls. We propose held-out evaluations for model changes, cross-border adaptation, and scale, including a factorial test of source evidence and checkpoint repair and group-level reporting to prevent aggregate gains from masking local failures. This is a methodological proposal; implementation and empirical validation remain future work.