Web agents improve safety by changing webpage state before acting
Before Acting, Change the State: Prospective State Intervention for Web Agents under Deceptive Interfaces
Artificial IntelligenceCryptography and Security
Summary
Large language model (LLM) based web agents can do tasks automatically on websites, but tricky or deceptive web pages can make them take actions that harm user interests. The authors found that sometimes an agent’s action is allowed but still leads to bad results because of the current webpage’s state. To fix this, they made a system called Veer that changes the webpage’s state first, guiding the agent to a safe setup before letting it act. Tests showed Veer works better than other defenses and stops deceptive tricks across many types of challenges.
What this means in practice
- •For web developers: Enhance web automation tools by adding state modification to prevent deceptive actions on tricky websites.
- •For automation engineers: Build safer autonomous web agents that preemptively adjust page state to avoid harmful or unauthorized outcomes.
Authors
Ruozhao Yang, Mingfei Cheng, Xiaofei Xie
Abstract
LLM-based Web agents can autonomously complete user tasks, yet deceptive interfaces can steer them toward outcomes that conflict with users' interests. Existing defenses primarily intervene on agent behavior through blocking, guidance, or replanning. We identify a distinct failure mode: a task-valid action can still realize an unauthorized consequence because of the current Web state. This motivates treating task-relevant Web state itself as a runtime control target. We introduce Veer, an agent-side runtime defense that leaves task planning to the base agent and intervenes on Web state when a proposed action would produce an unauthorized consequence. Before modifying the live environment, Veer constructs a prospective intervention trajectory toward a safe task-relevant state and executes it with runtime grounding and verification. Across TrickyArena and WebDecept, Veer achieves the highest safe task completion in all three evaluation settings, exceeding the next-best defense by 15.9 and 25.0 percentage points on TrickyArena-Single and TrickyArena-Multi, respectively, while reducing dark-pattern success on WebDecept to 0.3%. These gains persist across dark-pattern types and all 12 agent, model, and benchmark configurations. Ablations show that active state intervention provides the largest gain, while prospective rollout and temporal evidence contribute additional improvements. These results establish task-relevant Web state as an effective runtime control target for protecting Web agents from deceptive outcomes.