Motion based captchas reveal human perception edge over gui agents
Invisible in Space, Visible in Time: Motion Vision CAPTCHA against GUI Agents
Computer Vision and Pattern Recognition
Summary
Many captchas are solved by looking at static pictures, but smarter computer programs can now recognize and click through these easily. The authors created a new kind of captcha that relies on moving images, making it easy for humans but very hard for current computer agents to solve. This happens because the important parts of the image only appear clearly when watching how things move over time, rather than from a single snapshot. Their tests showed humans scored nearly perfect, while computer agents barely did better than guessing.
What this means in practice
- •For website security teams: Deploy captchas that use motion-based puzzles to prevent automated browser agents from passing authentication tests.
- •For interactive game developers: Incorporate motion-defined perception challenges that humans solve easily but current AI agents do not, improving user engagement.
Authors
Zeyu Zhang, Dingyi Rong, Zijian Chen, Zicheng Zhang, Xiongkuo Min, Guangtao Zhai
Abstract
Most existing visual CAPTCHAs remain spatially solvable: the required information is exposed by static appearance, local structure, and interface state. This assumption is weakened by advances in multimodal large language models (MLLMs) and Graphical User Interface (GUI) agents, which exhibit strong visual perception, reasoning, and browser interaction capabilities. We propose Motion Vision CAPTCHA (MVCAP), a hierarchical motion-based CAPTCHA framework in which target semantics are instantiated as motion-defined foreground structures and become recoverable only through temporal segregation from a dynamically evolving background. Built on this shared principle, MVCAP is instantiated in three perceptually progressive levels: coherent motion, structural motion, and biological motion. To evaluate this framework, we introduce MVCAP-Bench, a browser-based benchmark with 600 live CAPTCHA instances, together with a matched foreground-only control benchmark, MVCAP-Bench-FG. We evaluate humans, Browser Use agents, native computer use agents, and a supplementary offline VQA setting derived from the same instances. Results reveal a substantial human--agent gap: on the full MVCAP-Bench, human accuracy reaches 99.6%, whereas the best GUI agent achieves only 16.8%, close to the six-way chance level. The foreground-only control further shows that the key difficulty comes from dynamic background camouflage rather than answer format or browser interaction alone. These findings identify a measurable human--agent perception gap and position MVCAP-Bench as a benchmark for studying motion-defined perception in current agents.