Adaptive guard models interpret custom safety policies for language agents

AdaGuard: An Adaptive Guard Model with User-defined Policies

Artificial Intelligence

Summary

Language model agents must follow safety rules, but these rules can change a lot depending on where they are used. Fixed guard systems struggle with this because they can’t easily adjust to different rule sets. The authors created a new dataset and learning method to help models understand and check if an agent breaks customized safety rules. They trained a model called AdaGuard that can look at an agent’s actions and decide if they follow the rules given at the time, improving safety checks across many scenarios.

What this means in practice

  • For ai safety teams: Evaluate language model agents for compliance with custom safety policies during deployment to reduce risky or harmful outputs.
  • For ai platform engineers: Integrate adaptive guard models to dynamically interpret and enforce evolving safety rules on language model agent behaviors.

Authors

Yunhao Feng, Yifan Ding, Yuxiang Xie, Zheng Li, Mingrui Lao, Zeyuan Wang, Yanming Guo

Abstract

Guard models support the safe deployment of language model agents, but fixed risk taxonomies limit their ability to accommodate requirements that vary across applications and tasks. Under user-defined policies, detecting violations requires interpreting both the applicable rules and the agent's behavior, since identical actions can receive different judgments under different policies. To support learning this capability, we introduce AdaptiveSafety, a dataset of 10,939 training examples and 1,000 test examples covering policies with 1--100 rules. The dataset combines trajectories from multiple sources with policy and behavioral counterfactuals, pairing each example with an explanation and the complete set of violated rules. These counterfactuals expose changes that alter compliance, while structural augmentations provide supervision for consistency under rule reordering and identifier remapping. Building on this supervision, we propose SafePO, a reinforcement learning algorithm for refining violation identification while balancing explanatory reasoning and final verdicts. SafePO uses structured rewards to assess prediction correctness, retains group-relative advantages at the response level, and employs a separately trained value model to modulate token weights within explanation and verdict regions. Separate normalization controls their relative contribution to training despite differences in length. Through supervised initialization followed by SafePO, we develop AdaGuard, a family of 0.6B, 4B, and 8B guard models that assess agent trajectories under policies supplied at inference time. Our 4B model achieves binary accuracies of 89.30\% on AdaptiveSafety and 71.82\% on DynaBench. The project repository is available at https://github.com/Yunhao-Feng/AdaGuard