Papers for

software operations teams

Papers whose findings have a practical use for this group, as judged from the abstract. Open a paper to read what it means in practice.

JEV model tests agent security faster and cheaper than AI judges

JEV as a Judge for Agent Trace Security: An Empirical Comparison with Generative LLM Judges

Abstract: Security evaluation of tool-using agents requires judging actions in context, yet generative judges add latency, explanation overhead, and output-validation failures. We study whether JEV, a typed decision model, offers a useful alternative for retrospective trace classification. We evaluate JEV and four generative judges on four benchmark collections totaling 5,219 trajectories, using a common risk rubric and behavior-level labels. JEV attains a benchmark-averaged positive-class F1 of 77.8, compared with 74.1 for the strongest generative configuration, GLM-5.2, with valid-result coverage of 95.5\% and 94.4\%, respectively. Performance varies across datasets, with JEV leading on ATBench500 and MCPHunt and GLM leading on R-Judge and TraceSafe. Across the four benchmarks, JEV's median successful-call latency is 0.99 seconds; estimated token cost averages \$0.000195 per valid judgment. These results support JEV as an economical screening signal, with trade-offs in precision and recall.

Mon 28 SeptCryptography and Security
The gist
Security checks on the behavior of software agents usually rely on AI systems that generate explanations but can be slow and sometimes give unclear answers. The authors studied a tool called JEV, which uses a rule-based decision approach to check past actions for security risks. They found that JEV performs similarly to the best AI judge overall, is faster, and costs less while still covering most cases. This makes JEV useful as a fast initial filter for spotting security issues, although it may miss some details compared to AI methods.
Open → 2609.34862v1

Architecture enables AI agents to work on tasks lasting many days

An Architecture for Long-Horizon Agents: Levels, Ticks and Cascaded Intelligence

Abstract: Language-model agents are increasingly asked to carry out work spanning days or weeks, such as an operations remediation or a research programme. Such a task outlives any context window, any process and any interval at which a person can attend. In this paper, we argue that a long-horizon agent must run continually without forgetting before it can learn continually. This ability lies in the harness around the model rather than in the model itself. We derive seven bottlenecks from the long-horizon setting and answer them with a hierarchical architecture of three parts: (i) levels indexed by time scale, each keeping a bounded file summarising the level below; (ii) a clocked tick as the unit of autonomous action; and (iii) cascaded intelligence, where work is escalated to a more capable model only after failing review. We report on a ten-day campaign in which an agent built on this architecture reproduced a published reinforcement-learning result with a human attending once a day, and show (1) the agent kept the thread across every context reset and session boundary of the campaign, (2) operating knowledge written early changed later behaviour with no change to model weights, and (3) where learned components would enter such a system. Overall, our experience suggests continual learning for these agents needs a substrate outliving every context and process, and the checks the harness already runs are where a learner belongs.

Thu 17 SeptArtificial IntelligenceMachine Learning
The gist
Some AI helpers need to do jobs that take many days or weeks, but they forget too much and can’t keep track over time. The authors show that these helpers need a special system around the AI to remember and keep learning without losing old knowledge. They made a setup with three parts: different time levels that summarize what happened before, regular actions called ticks, and a way to pass problems up to smarter parts only if needed. They tested their system on a ten-day task where a human checked in just once a day, and it kept track and improved without changing the core AI.
Open → 2609.19519v1