Summary
Robots that learn to copy human movements can struggle when objects that look similar are nearby, causing them to grab or place items incorrectly. The authors studied how these robots get confused depending on what part of the action they are performing and the kind of visual similarity involved. They tested ways to help robots focus on the right target by making them pay more attention to the important object and teaching them to ignore lookalikes. These fixes helped robots do better in both computer simulations and real-world tests, even in complicated tasks like handling medical tools. This work shows that helping robots better recognize the right object can make their actions more reliable.
visuomotor imitationvisual groundingdistractor objectsaction chunkingtransformersrobot manipulationattention regularizationvisual promptingrobot learningrobot control
Authors
Vivek Chavan, Pengtao Xie, Yahuan Shi, Oliver Heimann, Kevin Haninger, Jörg Krüger
Abstract
Visuomotor imitation policies can achieve high performance under in-distribution visual conditions yet fail when visually similar objects or receptacles are introduced. We study this behavior as a problem of conditional visual grounding: the visual target required for successful control changes with the manipulation phase and, in more complex tasks, with the observed task state. Using Action Chunking with Transformers (ACT), we systematically introduce distractor objects and receptacles with controlled color and shape similarity and localize failures to picking and placement. We find that distractor sensitivity is specific to both the type of visual similarity and the manipulation stage. Guided by this diagnosis, we evaluate distractor augmentation, phase-dependent attention regularization, and appearance-based visual prompting as complementary interventions for improving target selection while preserving spatial information required for control. These interventions substantially improve robustness in simulation and on a physical UR3e. We further examine the same failure pattern in a pretrained vision-language-action policy on a state-conditioned instrument-handling task, where the observed state of a medical instrument determines the correct destination. Together, the results show that visual distractors can cause incorrect object or destination selection even when the underlying manipulation skill remains intact, and that explicitly improving target selection can substantially recover performance across distinct visuomotor policy-learning regimes.