JEV model tests agent security faster and cheaper than AI judges

JEV as a Judge for Agent Trace Security: An Empirical Comparison with Generative LLM Judges

Cryptography and Security

Summary

Security checks on the behavior of software agents usually rely on AI systems that generate explanations but can be slow and sometimes give unclear answers. The authors studied a tool called JEV, which uses a rule-based decision approach to check past actions for security risks. They found that JEV performs similarly to the best AI judge overall, is faster, and costs less while still covering most cases. This makes JEV useful as a fast initial filter for spotting security issues, although it may miss some details compared to AI methods.

What this means in practice

  • For security engineering teams: Screen automated software behaviors for security risks rapidly and cost-effectively using a typed decision model as a first pass.
  • For software operations teams: Reduce overhead of security trace analysis by employing a low-latency classification method that requires less output validation.

Authors

Zhiqiang Wang, Yichao Gao

Abstract

Security evaluation of tool-using agents requires judging actions in context, yet generative judges add latency, explanation overhead, and output-validation failures. We study whether JEV, a typed decision model, offers a useful alternative for retrospective trace classification. We evaluate JEV and four generative judges on four benchmark collections totaling 5,219 trajectories, using a common risk rubric and behavior-level labels. JEV attains a benchmark-averaged positive-class F1 of 77.8, compared with 74.1 for the strongest generative configuration, GLM-5.2, with valid-result coverage of 95.5\% and 94.4\%, respectively. Performance varies across datasets, with JEV leading on ATBench500 and MCPHunt and GLM leading on R-Judge and TraceSafe. Across the four benchmarks, JEV's median successful-call latency is 0.99 seconds; estimated token cost averages \$0.000195 per valid judgment. These results support JEV as an economical screening signal, with trade-offs in precision and recall.