Papers for

web automation teams

Papers whose findings have a practical use for this group, as judged from the abstract. Open a paper to read what it means in practice.

Gui agents learn from one try to improve task success rates

One Rollout Is All You Get: Fully Test-Time Adaptation for GUI Agents

Abstract: GUI agents are deployed with frozen weights and discard everything they experience on the job. Existing ways to update an agent's weights assume something deployment withholds: ground truth, rollouts beyond the single attempt (retries, samples, practice runs), or a learning phase other than deployment. Because GUI actions can be irreversible, a deployed agent gets one attempt per task occurrence, in arrival order, and every attempt counts. No ground truth is available at any point. We define fully test-time adaptation for GUI agents by these constraints and pair it with a minimal weight-space method, SOLO. Auxiliary models read each episode: a judge selects the episodes it deems successful, and a proposer-verifier pair relabels a failed episode's prefix with the subtask that prefix completed. Admitted episodes enter a short sliding window, and each admission updates a small adapter by top-K self-distillation on the agent's own predictions, provided the window holds a judged success. On recurring task streams built from WebArena, VisualWebArena and MobileWorld, SOLO improves on the frozen agent with both UI-TARS-7B and Qwen3-VL-8B, by three to six points of success rate, and exceeds two in-setting memory methods on the web streams.

Mon 28 SeptMachine Learning
The gist
Graphical user interface (GUI) agents usually can’t learn after being put to work because they never get feedback or multiple chances. The authors designed a way for these agents to adapt during their single chance at each task, even when no answers or retries are available. Their method, called SOLO, uses internal checks to figure out if an attempt was good or where it went wrong, then updates the agent based on recent successful tries. This approach helped agents perform better on various web and mobile task challenges.
Open → 2609.34321v1

Structured diagnosis uncovers key failure points in GUI agents

GUITAR: Structured Failure Diagnosis of GUI Agents via State Transitions

Abstract: Understanding where and why Graphical User Interface (GUI) agents fail is essential for building more reliable systems, yet current evaluation relies on step accuracy, a metric that treats each screen independently and overlooks the underlying structure of GUI environments. This leads to two critical blind spots: (1) functionally equivalent screens are evaluated in isolation, obscuring systematic failure patterns across shared screens; and (2) the long-tailed GUI distribution renders failures on rare but critical screens invisible under standard metrics. To address these issues, we propose \textbf{GUITAR}, a state-centric diagnostic framework that performs structured failure analysis over both states and transitions, using a State Transition Graph (STG) by mapping visually diverse screens to shared functional states. Across 8 agents and 6 tasks from AndroidControl and Mind2Web, GUITAR reveals that 60.4\% of failures occur in 20\% of states, localizing errors to a small set of bottlenecks. Bottleneck-targeted guidance improves SR by 2.8\% and retains a 1.88\% average gain across 7 agents under three-fold trajectory-held-out evaluation with fully automatic STGs. These findings demonstrate the diagnostic and actionable value of structure-aware evaluation within the evaluated mobile and web tasks. Code is available at https://github.com/sqzhang-lazy/GUITAR

Mon 28 SeptArtificial Intelligence
The gist
Graphical User Interface (GUI) agents sometimes fail in ways that are hard to spot because traditional evaluations treat each screen separately. The authors propose GUITAR, a method that groups visually different screens into shared states and studies how agents move between these states. This approach uncovers that most failures happen in a small number of key screens, which can then be targeted for improvement. By focusing on these bottlenecks, agents perform better on tasks across various mobile and web environments.
Open → 2609.34113v1