Structured diagnosis uncovers key failure points in GUI agents
GUITAR: Structured Failure Diagnosis of GUI Agents via State Transitions
Artificial Intelligence
Summary
Graphical User Interface (GUI) agents sometimes fail in ways that are hard to spot because traditional evaluations treat each screen separately. The authors propose GUITAR, a method that groups visually different screens into shared states and studies how agents move between these states. This approach uncovers that most failures happen in a small number of key screens, which can then be targeted for improvement. By focusing on these bottlenecks, agents perform better on tasks across various mobile and web environments.
What this means in practice
- •For mobile app developers: Identify critical app screens where automated GUI agents frequently fail and focus testing efforts for more reliable automation.
- •For web automation teams: Improve failure diagnosis of web UI automation by analyzing functional states rather than isolated pages, leading to targeted fix strategies.
Authors
Shaoqing Zhang, Kehai Chen, Xuefeng Bai, Zhuosheng Zhang, Pengfei Zhang, Yang Xiang, Min Zhang
Abstract
Understanding where and why Graphical User Interface (GUI) agents fail is essential for building more reliable systems, yet current evaluation relies on step accuracy, a metric that treats each screen independently and overlooks the underlying structure of GUI environments. This leads to two critical blind spots: (1) functionally equivalent screens are evaluated in isolation, obscuring systematic failure patterns across shared screens; and (2) the long-tailed GUI distribution renders failures on rare but critical screens invisible under standard metrics. To address these issues, we propose \textbf{GUITAR}, a state-centric diagnostic framework that performs structured failure analysis over both states and transitions, using a State Transition Graph (STG) by mapping visually diverse screens to shared functional states. Across 8 agents and 6 tasks from AndroidControl and Mind2Web, GUITAR reveals that 60.4\% of failures occur in 20\% of states, localizing errors to a small set of bottlenecks. Bottleneck-targeted guidance improves SR by 2.8\% and retains a 1.88\% average gain across 7 agents under three-fold trajectory-held-out evaluation with fully automatic STGs. These findings demonstrate the diagnostic and actionable value of structure-aware evaluation within the evaluated mobile and web tasks. Code is available at https://github.com/sqzhang-lazy/GUITAR