WebPageBench verifies web agents by matching event logs across UI variants
WebPageBench: Event-Level Verification and Controlled UI-Variant Generation for Web Agents
Artificial IntelligenceComputation and Language
Summary
Many computer programs called web agents act on websites to do tasks, but testing if they really succeed is hard. The researchers built WebPageBench, a tool that tracks every action on test websites by logging the exact events triggered during the task. This helps check if agents truly complete tasks without just scraping pages. WebPageBench can also switch website controls’ designs to see how changes affect agent performance without changing the task itself.
What this means in practice
- •For web engineers: Evaluate web automation tools by verifying task success from user interface event logs precisely.
- •For software testers: Measure how changes in website controls affect automated agent success under fixed task conditions.
Authors
Anton Emelyanov, Maria Tikhonova, Zaven Martirosian, Sergei Averkiev, Alena Fenogenova
Abstract
We present WebPageBench, an open framework for evaluating web agents in which every task is verified from the interface's own event log. Six instrumented mock sites with brand identifiers removed (a marketplace, a bookstore, a grocery service, rail ticketing, hotel search and a document cabinet) emit typed events with parameters as a user or an agent acts. A task declares the events it requires, and success is decided by matching them, with no judge model and no scraping of rendered pages. The same instrumentation supports controlled UI variation: one configuration switch re-renders a task through a different implementation of a single control while the prompt and the success conditions stay completely identical, so sensitivity to interface form can be measured under a fixed task specification. The WebPageBench release consists of three components: 152 tasks, divided into 65 canonical scenarios and 87 control variants across light/dark UI-modes; a common runner evaluated with six browser/DOM harness configurations and five screenshot-only GUI-agent families; and a public leaderboard of 24 model-harness pairs. On the public 152-task leaderboard the gap between what agents declare finished and what the log confirms reaches 41 points (one configuration declares every task finished and satisfies the conditions on 59%).