Summary
Robots that act based on visual or language instructions often struggle to understand the exact positions and shapes of objects around them. This paper introduces a method called Spatial Grafting, which links 3D shape information directly to the robot’s viewpoint without changing the robot’s original way of seeing the world. The authors show that adding this 3D-grounded spatial information helps robots perform better on a variety of tasks, including moving objects with one or two arms and working in cluttered or changing environments, both in simulation and on real robots. Their approach improves performance more than previous methods and works across different robot systems and tasks.
What this means in practice
- •For robotics engineers: Build manipulation controllers that use 3D spatial features grounded in robot coordinates to improve grasping and moving objects under different conditions without retraining perception.
- •For industrial automation teams: Enhance dual-arm robot performance in complex tasks like cluttered pick-and-place and mobile manipulation by integrating spatial grafting techniques for better environment understanding.
Authors
Dingsheng Liu, Yangzheng Wu, Mahboubeh Asadi, Zhiyuan Li, Jinbang Huang, Yixin Xiao, Tongtong Cao, Yingxue Zhang
Abstract
Pretrained robot manipulation policies such as vision-language-action models (VLAs) or world-action models (WAMs) leave interaction-relevant metric geometry implicit. Recent breakthroughs in spatial reconstruction can supply the necessary geometry reliably, but their features describe local shape without stating where it lies with respect to the robot. How best to deliver these features to a pretrained policy remains unresolved. We propose Spatial Grafting, a versatile, lightweight spatial module that binds frozen reconstruction features to metric, robot-relative geometry. Spatial Grafting constructs metric-grounded spatial tokens and injects them into the flow-matching action expert through cross-attention, without modifying the host's perceptual pathway, so the host retains the full benefit of its pretraining. We evaluate it more broadly than any geometry-aware policy we compare against: one graft architecture, with no per-host redesign, on two VLAs and two WAMs, across four simulation benchmarks that span short-horizon manipulation, visual robustness, clutter and long-horizon mobile manipulation, and on three real-robot platforms with single- and dual-arm configurations. On RoboTwin 2.0, a dual-arm manipulation benchmark, the graft improves every host across VLAs and WAMs. Grafted $π_{0.5}$ gains 11.3% and 15.6% on clean and randomized scenes, reaching 94.0% and 92.4%, above the strongest published 3D-conditioned policy, WAM4D (93.8% and 89.9%). The margin widens as the horizon lengthens: on tasks from BEHAVIOR-1K, a dual-arm mobile manipulation challenge scored by average task progress, it surpasses the 2025 challenge winner on five of six tasks,by up to 0.47 Q-score, and exceeds a map-conditioned spatial policy on average across the three tasks both report.