Failure-transparent agents reduce false success claims in AI tools
Failure-Transparent Agents: Benchmarking Post-Failure Reporting in Tool-Using Language Models
Artificial Intelligence
Summary
Sometimes AI agents use tools that fail, and then the agents incorrectly say they succeeded without proper proof. The authors created a special test called Failure-Transparent Agents (FTA) to focus on checking whether agents honestly report failure after tools fail. They tried different models and methods, showing that adding clear instructions and requiring structured evidence drastically reduces false success claims and made responses more helpful. This work helps improve trustworthiness in AI that uses other tools.
What this means in practice
- •For ai system developers: Implement structured evidence contracts in AI agents to reduce false success claims and improve reliability after tool failures.
- •For customer support automation teams: Use failure-transparent agents to ensure automated responses accurately reflect when backend tools or data sources have failed.
Authors
Junru Zhu, Shiming Xie, Aime Lu Fan Chen, Xiaoqing Ding, Chunxin Tang, Ruoyu Qi, Yulang Fei
Abstract
Tool-using agents can fail twice: a required tool can fail, and the agent can then report success without the evidence needed to justify it. Existing benchmarks often entangle this reporting failure with tool selection, recovery, and environment dynamics. We introduce Failure-Transparent Agents (FTA), a controlled benchmark that fixes the failed observation and required evidence state before generation, making post-failure claims directly auditable. FTA contains 100 tasks with deterministic failure traces spanning five failure families, a neutral control, and four user-pressure conditions, and evaluates unsupported claims alongside useful recovery. Across six models, three response policies, and 3,600 human-annotated responses, false-success rates are 22.8% under the baseline policy, 9.3% with a transparency instruction, and 0.8% with a structured evidence contract. Fabricated-detail rates decrease from 28.3% to 14.3% and 0.8%, while useful responses increase from 74.9% to 89.2% and 98.8%, respectively. The tested evidence-contract policy is associated with substantially lower post-failure reporting errors while useful-response rates remain high within this blocked-task benchmark.