Future duet improves robot control by separating camera views
FutureDuet: Decoupling Observation Access from Future Supervision in World Action Models
Robotics
Summary
Robots often use cameras to watch and learn how to do tasks, but different cameras see different angles with different details. The authors found that treating the main camera and wrist camera separately with different prediction goals helps the robot understand both the whole scene and close-up actions better. Their approach, called FutureDuet, improves robot success on difficult tasks that require precise interactions. The method adds no extra work during use, only during training.
What this means in practice
- •For robotics engineers: Enhance robotic manipulation by separately processing fixed and moving camera views to improve task success rates involving precise interactions.
- •For industrial automation teams: Deploy robots in complex assembly lines with improved action prediction from future visual information of both scene-level and close-range actions.
Authors
Jie Wu, Yuzhi Huang, Junqi Liu, Weichen Zhang, Haibin Huang, Yin Chen, Jingyan Jiang, Chi Zhang
Abstract
World Action Models (WAMs) augment robot action generation with future visual supervision. Existing WAMs commonly fuse main and wrist observations into one visual stream and train both with the same future-video objective, despite their different visual dynamics. A stable main camera reveals scene-level task evolution, whereas wrist cameras move with the end effector, mixing local interaction changes with viewpoint shifts and self-occlusion. These contrasting predictive demands suggest that the two views may benefit from different future objectives. We introduce FutureDuet, which retains both views for control, while allowing each visual stream to receive a different future objective. For the main view, future RGB models task evolution, while interaction masks and robot skeletons focus supervision on task objects and robot motion. For the wrist stream, future latent prediction models short-horizon interaction changes without requiring pixel-level reconstruction. ActionDiT jointly reads the resulting Task State and Interaction State, combining scene-level progress with close-range interaction evidence. All auxiliary prediction modules are training-only, adding no inference overhead. FutureDuet achieves 94.2% clean and 94.1% randomized success on RoboTwin50 and 99.2% average success on LIBERO. The improvements are most pronounced on six RoboTwin50 tasks that require precise interaction, averaging gains of 9.2% and 12.8% over Fast-WAM in clean and randomized settings. Controlled studies further show complementary gains from separating the wrist pathway and designing future supervision separately for the two views.