Gaze prompts improve robot manipulation success using eye tracker data
Gaze Prompts: Temporally Dense Human Attention for Vision-Language-Action Fine-Tuning
Robotics
Summary
Robots often get instructions at a high level but don’t know exactly what to pay attention to in each moment of a task. This paper shows how using eye-tracking data from humans controlling robots in virtual reality can help guide the robot’s attention every step of the way. The researchers turned recorded eye gaze into visual cues that train the robot to focus better on important parts of the scene. This approach nearly doubled the robot’s success on tricky two-handed tasks.
What this means in practice
- •For robotics engineers: Create robot control systems that use human gaze data to improve real-world manipulation task success.
- •For virtual reality developers: Build teleoperation interfaces that record operator gaze to enhance robot training and task performance.
Authors
Yihan Zhou, Rui Yan, Mingcong Li, Zheyuan Huang, Xu Yang, Xueyang Guo, Yilin Mo
Abstract
Vision-Language-Action (VLA) fine-tuning pairs images with actions at every step, yet typically provides only a task-level language instruction, leaving moment-to-moment visual relevance implicit. We introduce \emph{eye-tracker-supervised gaze prompting}, which uses gaze recorded during VR teleoperation to provide frame-level visual guidance for VLA fine-tuning. During training, recorded gaze locations are rendered as crosshairs on the robot's head-camera images. At deployment, a lightweight predictor estimates gaze locations from recent images and the instruction, supplying the same type of visual prompt without an eye tracker or changes to the policy architecture. Instantiated with $π_0$, gaze prompting increases mean success from $26.3\%$ to $56.0\%$ across six real-world bimanual manipulation tasks, with gains also observed when a single policy is trained on all six tasks. We release \textsc{GazeMani}, a dataset of $1{,}200$ teleoperated trajectories with synchronized gaze.