Aggregate agent behavior diagnosability improves fault detection accuracy

When Are Aggregate Agent Traces Diagnosable? Traffic-Governed Interpretation and Calibrated Abstention

Artificial Intelligence

Summary

When computer agents choose actions based on complex rules, some problems or faults might not be obvious because the agent rarely visits the affected situations. The authors study a pricing agent simulation for hotels to find when faults show clearly in the aggregate behavior and when they do not. They propose a way to decide if the observed data is sufficient to diagnose a fault or if the system should abstain from making a claim. Their method reduces false alarms and improves confidence in detecting real problems by ensuring the agent has enough exposure to the suspected faulty conditions before interpreting the traces.

What this means in practice

  • For runtime system engineers: Identify when faults in agent-driven systems can be confidently detected based on observable behaviors and abstain otherwise to reduce false detections.
  • For automated pricing teams: Improve fault diagnosability in dynamic pricing systems by establishing conditions for valid interpretation of aggregate agent actions under varying market demands.

Authors

Peiying Zhu, Sidi Chang

Abstract

Runtime traces can appear transparent, but a closed-loop policy determines which states are visited and which failures become visible. We study a simulated hotel-pricing agent mapping time, inventory, and market state to discrete price actions under varying demand regimes. A fault may leave no aggregate trace when the policy rarely visits affected cells. We treat entry into aggregate-only fault interpretation as a diagnosability decision preceding scoring or localization. A reference-map gate requires repeated clean-policy support; a matched runtime gate then requires joint support in clean and current streams. Signal analysis occurs only after both pass. We calibrate false admission on a disjoint clean stream at the physical-component level and model detection by affected clean traffic rather than nominal cell coverage. In a frozen one-shot heldout, 55/72 (76.4%) regime-component units were reference-admitted, representing 20 physical components; 54/55 passed matched runtime admission, while the rejected unit abstained. Stable false admission was 0/20, with a one-sided exact 95% upper bound of 0.1391, meeting the frozen 0.20 criterion. Across 540 repeated unit-arm rows nested in those 20 clusters, affected clean traffic reduced negative log likelihood by 29.3% relative to cell coverage, a gain of 0.1264 nats per row (cluster-bootstrap 95% interval [0.0593, 0.1918]). Adding mask family and its interaction improved log loss by 0.0015 nats per row (one-sided upper bound 0.0066), below the frozen 0.01 practical-sufficiency margin. A development audit found that exact minimum hitting set and greedy selection chose identical supports in 12/12 scenarios because singleton evidence had resolved the conflicts. The result is a bounded rule for interpreting aggregate agent behavior: first establish exposure, then score change, and abstain when the trace cannot support the claim.