Safety system cuts risky behaviors in real-time AI agents
PROACT-Agent: Progressive Runtime Oversight and Active Circuit-breaking for Real-Time Safety
Machine Learning
Summary
Large language models used as AI agents can cause harmful effects beyond just bad text, like damaging the environment. Current safety checks happen too late, after harm is done. The authors propose PROACT-Agent, a new way to watch and intervene during AI actions to prevent risks before they happen. They created new datasets and methods to catch hidden dangers and keep AI behaviors safe across different cultures. Their tests show their system can accurately spot unsafe actions and drastically reduce successful attacks.
What this means in practice
- •For ai system developers: Integrate PROACT-Agent to detect and prevent unsafe actions in AI agents during operation, improving real-time safety in complex tasks.
- •For cybersecurity teams: Use the approach from PROACT-Agent to reduce the success rate of attacks targeting AI systems within operational environments.
Authors
Ding Jia, Wei Liu, Xianglong Du, Yingjie Li, Yingqing Yang, Huili Yu, Zhangsong Zhan, Chu Zhou
Abstract
The transition from Large Language Models (LLMs) to agents shifts safety stakes from toxic text to irreversible environmental harm. While current defenses remain largely retrospective, proactive runtime intervention is bottlenecked by the lack of large-scale, causally-consistent data. We propose PROACT-Agent, a framework for synthesizing high-fidelity trajectories to enable real-time guardrails. We identify a critical "safety drift" in prior benchmarks, where lenient annotation paradigms fail to enforce temporal consistency. PROACT-Agent addresses this through: (1) Progressive Trajectory Unrolling to reveal risks hidden in long-context interactions; (2) Reasoning-Augmented Causal Rectification to enforce monotonic causal consistency; and (3) Culturally-Aware Data Localization for cross-border robustness. We introduce PROACT-Bench, a bilingual safety benchmark with 155,780 states labeled through multi-model adjudication. Evaluating updated context before the next LLM inference, the trained guard achieves 91.46% unsafe-class F1 and 90.63% exact-boundary detection under complete source holdout. In AgentDojo, it reduces non-DoS targeted attack success from 20.82% to 0.40%.