Gaze-driven head movement reveals passcodes from afar
GAZEleak: Passcode Inference Against Eye-tracking XR Devices Through External Observation
Cryptography and Security
Summary
Mixed-reality headsets let users input passcodes by looking and pinching, which was thought to be safe from people watching. This paper shows that just filming the wearer from a distance can reveal passcodes by detecting tiny head movements that happen when their eyes shift gaze. The authors’ method guesses codes much better than random chance without needing anything inside the headset. This finding suggests a new kind of privacy risk for eye-tracking devices.
What this means in practice
- •For security teams: Identify physical observation risks for gaze-based input devices to improve user privacy defenses.
- •For device manufacturers: Design hardware or software mitigations to reduce head-movement leakage of gaze-driven input in mixed-reality headsets.
Tested on one dataset.
Authors
Hwanjo Heo, Junhee Lee, Jinwoo Kim
Abstract
Mixed-reality headsets such as Apple Vision Pro replace the touch screen with gaze-as-pointer interaction: the wearer looks at a target and confirms with an air pinch. Because the display is inside the headset and the eye tracker is walled off from third-party software, such input is widely assumed to be unobservable to bystanders---a built-in defense against the shoulder-surfing that plagues phones and laptops. We present GAZEleak, a side-channel attack that recovers gaze-driven input from external video of head motion alone, under a strictly local, physical-observer threat model: the adversary only films the wearer from across the room and installs no software on the device. The attack exploits the centrally coupled eye-head motor program: gaze shifts recruit small, target-dependent head reorientations that project into sub-degree pose changes recoverable from commodity video. GAZEleak implements a measurement-based, sparse-optical-flow inference pipeline for users' 6-digit device passcodes. On a preliminary front-view dataset from three author-subjects who were aware of the attack hypothesis, GAZEleak places the true code within the top ten guesses for 56% (10 of 18) of test codes under a cross-person protocol with no labeled victim data, and for every code of the most exposed subject. The performance is subject-dependent: no passcode from the least exposed subject reaches the top ten, although its median guessed passcode rank is 12,786 rather than 500,000 expected from an uninformative ordering. These results provide preliminary evidence that gaze-coupled head motion can expose passcode information under controlled conditions, while motivating broader evaluation across users, behaviors, and capture settings.