Papers for

ai system reliability teams

Papers whose findings have a practical use for this group, as judged from the abstract. Open a paper to read what it means in practice.

Reliability engineering methods adapted to ensure ai system safety

Reliability Engineering for AI Systems: Challenges, Methods, and Directions

Abstract: AI reliability concerns whether an AI system performs its intended function dependably over a stated period and under stated operating conditions, with stated evidence. As these systems become more autonomous, that function includes more than a correct output. Retrieval, memory, tool use, permissions, human oversight, and interactions among systems must operate consistently and safely, and, for generative systems, so must the reasoning process that produces the output. Average benchmark accuracy measures capability; it does not quantify this broader reliability claim. This paper adapts established reliability engineering methods, from failure definitions and operational envelopes to FMEA, accelerated testing, field monitoring, and reliability growth, to AI systems. A four-level diagnostic framework classifies failures as component, operational-loop, agentic-conduct, or network and governance failures. Test, evaluation, verification, and validation (TEVV), sequential monitoring, and FRACAS create and refresh evidence. SMART provides statistical guidance for measurement, analysis, assessment, and test planning; the NIST AI Risk Management Framework provides organizational guidance for governance, evaluation, monitoring, and mitigation. Three cases illustrate the program: adversarial testing of a convolutional neural network, perception-error propagation, and autonomous-vehicle disengagements. Established reliability engineering provides a usable foundation; new measurements and safety guardrails are still needed as these systems are self-evolving.

Mon 28 SeptArtificial Intelligence
The gist
AI systems need to work reliably not just by giving correct answers but by functioning safely and consistently in many ways, such as using tools, remembering information, and interacting with people. The authors explain how traditional reliability engineering methods can be adapted to check and improve these complex AI behaviors. They introduce ways to detect different kinds of failures, continuously monitor AI performance, and support decisions with statistical guidance. Their approach is illustrated with examples like testing neural networks against tricky inputs and analyzing self-driving car safety. The authors emphasize that while established methods provide a useful base, new safety measures are necessary as AI systems grow more autonomous and evolve on their own.
Open → 2609.35316v1