Underwater predictive model helps robots manipulate objects without contact sensors
Underwater C3-JEPA: An Object-Centric Cross-View World Model for ROV Salvage
RoboticsArtificial Intelligence
Summary
Handling heavy objects underwater using remote-operated robots is hard because the robots can't feel what they touch and water slows their movements. The authors developed a method that uses multiple camera views and control commands to predict how objects will move when the robot interacts with them. This prediction happens in a compact form that focuses on the objects involved, making it efficient and accurate. Their approach works both in simulations and real underwater videos, helping robots plan better actions without needing physical contact sensors.
What this means in practice
- •For underwater robotics teams: Enable underwater vehicles to predict object states ahead of time without contact sensors using multi-view video and control inputs for better salvage operations.
- •For industrial automation engineers: Improve manipulation of heavy, visually complex objects in cluttered environments by applying multi-view control-conditioned predictive models beyond underwater scenarios.
Authors
Yuncong Yang, Jinlong Li, Yulong Xue, Feng Wu, Chunwen Zhang, Lei Qiao, Xuyang Wang
Abstract
We present Underwater C$^{3}$-JEPA (cross-view, control-conditioned, context-extended), an object-centric multi-view predictive world model for near-field heavy-load underwater ROV salvage. Without contact sensors, it predicts in latent space how the task-object state evolves through contact interaction and under the hydrodynamic lag of the vehicle, from synchronized multi-view RGB observations and vehicle control signals. C$^{3}$-JEPA encodes multi-camera observations into task-object and context tokens, fuses cross-camera evidence through held-out-view attention, and directly predicts future states conditioned on control. Weak binding anchors the target and gripper at low annotation cost, while SIGReg sharpens the geometric representation. Experiments show that the learned representation transfers substantially more task-relevant information to downstream probes than a reconstruction-free latent baseline, while keeping the predictor lightweight. The resulting predictive interface supports model-predictive-control (MPC) candidate evaluation and imagined-rollout behavior-agent training. Validation on real underwater video shows the same architecture recovering a withheld camera's object state and staying ahead of persistence, so the recipe transfers beyond simulation.