Test-Time Self-Evolving GUI Visual Grounding via Reflection-Guided On-Policy Self-Distillation

2026-08-11Computer Vision and Pattern Recognition

Computer Vision and Pattern RecognitionArtificial IntelligenceComputation and Language
AI summary

The authors study how to make computer programs better at understanding where to click or look in a graphical user interface (GUI) after they have already been set up and used. They created a system that lets the program try out tasks, check how well it did using a special language model, and learn from its own mistakes without needing new human help. This learning happens continuously while the program is running by turning feedback into detailed guidance for the model itself. Their experiments show this approach improves performance on several tests by a notable margin.

GUI Visual GroundingTest-Time AdaptationSelf-DistillationReinforcement LearningMulti-Lingual Large Language Model (MLLM)Self-Evolving FrameworkContrastive CalibrationOn-Policy LearningExploration-Reflection LoopAuto-Regressive Models
Authors
Shiyu Xuan, Zechao Li
Abstract
GUI Visual Grounding is a fundamental capability for GUI agents. Existing models typically freeze their parameters after deployment, limiting their ability to adapt to unseen interfaces. Although recent methods attempt to adapt models via test-time reinforcement learning, they cannot reflect upon failed exploration. To overcome this, we propose a Test-Time Self-Evolving framework that enables models to improve after deployment without human-annotated ground truth. It constructs a closed-loop of Exploration, Evaluation, Reflection, and Internalization. Specifically, the agent first explores unseen interfaces by predicting grounding coordinates for given instructions. To evaluate these explorations, we introduce an MLLM-based Reflector to assess the generated results and provide the corresponding reasoning reflections. To internalize reflection knowledge into the model weights, we propose Reflection-Guided On-Policy Self-Distillation, which translates high-level reasoning into dense token-level supervision via a conditioned self-teacher. Furthermore, we design a Contrastive Calibration method to prevent incorrect auto-regressive prefixes from corrupting the supervisory signals during failed explorations. Extensive experiments across six benchmarks demonstrate our framework's effectiveness, achieving an average accuracy improvement of 7.4% over the base model. To the best of our knowledge, this is the first work to successfully exploit on-policy self-distillation for test-time adaptation in GUI visual grounding. By filling the gap in post-deployment adaptation, our framework completes the self-evolving capability of GUI agents. The code will be released.