WebWorld: The Browser as a World Model for Self-Improving Web Code

2026-08-31Computation and Language

Computation and LanguageSoftware Engineering
AI summary

The authors identify a problem where AI models both suggest and judge fixes for web code, which doesn’t always ensure the code truly works. They introduce WebWorld, a system where a browser acts like a trustworthy judge by running the code and only approving changes that improve it without breaking anything. This process helps the AI learn better ways to generate web code. Their experiments show WebWorld leads to significant improvements compared to previous models, especially because it uses browser feedback to accept changes. Without this browser-based approval, the improvements mostly disappear.

Vision-Language Model (VLM)Web code repairBrowser simulationHTMLInteractive programmingSelf-improvement loopSoftware fine-tuning (SFT)Quality certificationAutomated code evaluation
Authors
Jiajun Wu, Jian Yang, Yaxin Du, Wei Zhang, Haowen Wang, Junhang Cheng, Yuxuan Zhang, Tuney Zheng, Xianglong Liu, Ming Zhou
Abstract
VLM-driven self-improvement of web code has a structural flaw: the model that proposes the repair is the model that judges it, and visual plausibility under that judge is a poor proxy for whether the page actually works. What the loop is missing is a counterparty the VLM cannot fool, and the browser already is that counterparty: a deterministic, executable simulator of how an HTML artifact behaves under user actions, and in everything but name a world model for web code. We present WebWorld, the interface that lets a VLM prior interact with this browser-as-world-model autonomously and decides which interactions become supervision. Each round, the VLM emits a critique that the planner compiles into a typed interaction contract; the browser re-executes the candidate and issues an acceptance certificate only when both target progress and preservation of every previously verified capability hold; certified transitions accumulate as a quality ratchet that is the only thing the SFT export ever sees. Under matched training, WebWorld-27B improves Raw-27B by 5.3 points on HTMLBench-400 and 14.9 points on MiniAppBench-Val, and reaches the level of strong frontier systems such as Kimi-K2.6 and GPT-5.4 on interactive HTML generation. Equal-size ablations show that browser-backed admission carries the gain: without the certificate, the matched 9B lift nearly disappears.