Gui agents learn from one try to improve task success rates

One Rollout Is All You Get: Fully Test-Time Adaptation for GUI Agents

Machine Learning

Summary

Graphical user interface (GUI) agents usually can’t learn after being put to work because they never get feedback or multiple chances. The authors designed a way for these agents to adapt during their single chance at each task, even when no answers or retries are available. Their method, called SOLO, uses internal checks to figure out if an attempt was good or where it went wrong, then updates the agent based on recent successful tries. This approach helped agents perform better on various web and mobile task challenges.

What this means in practice

  • For mobile app developers: Improve app agents’ ability to adapt and succeed on first user task attempt without needing ground truth or retries.
  • For web automation teams: Enhance web interface automation tools to update their behavior dynamically from live usage episodes without offline retraining.

Authors

Ziqiang Wang, Li Gu, Zhixiang Chi, Linlian Jiang, Zihuan Jiang, Linqiang Guo, Siobhan Reid, Zhi Liu, Yang Wang

Abstract

GUI agents are deployed with frozen weights and discard everything they experience on the job. Existing ways to update an agent's weights assume something deployment withholds: ground truth, rollouts beyond the single attempt (retries, samples, practice runs), or a learning phase other than deployment. Because GUI actions can be irreversible, a deployed agent gets one attempt per task occurrence, in arrival order, and every attempt counts. No ground truth is available at any point. We define fully test-time adaptation for GUI agents by these constraints and pair it with a minimal weight-space method, SOLO. Auxiliary models read each episode: a judge selects the episodes it deems successful, and a proposer-verifier pair relabels a failed episode's prefix with the subtask that prefix completed. Admitted episodes enter a short sliding window, and each admission updates a small adapter by top-K self-distillation on the agent's own predictions, provided the window holds a judged success. On recurring task streams built from WebArena, VisualWebArena and MobileWorld, SOLO improves on the frozen agent with both UI-TARS-7B and Qwen3-VL-8B, by three to six points of success rate, and exceeds two in-setting memory methods on the web streams.