Feed-forward method builds consistent 4D robot world models from multiple cameras

RoGSW4RLD: Feed-Forward 4D Gaussian Lifting for Robot World Model Rollouts

Computer Vision and Pattern RecognitionRobotics

Summary

Robots often use several cameras to understand and predict what will happen in their environment, but the video predictions they make are usually separate and hard to combine. The authors introduce RoGSW4RLD, a method that merges these multiple camera views into one consistent and interactive 3D space that changes over time. This approach uses the robot's movements and geometry to better combine different views and refine the scene, resulting in more accurate and coherent predictions. Tests show that their method gives clearer images and more precise depth and robot position estimates than previous techniques.

What this means in practice

  • For robotics engineers: Create unified 4D models from multi-camera video predictions to improve robot perception and planning accuracy.
  • For augmented reality developers: Generate consistent, time-varying 3D environments from multiple camera views for immersive AR experiences around moving robots.

Authors

Jin Hyun Kim, Min Young Kim, Soohwan Song, Daekyum Kim

Abstract

Action-conditioned video world models predict future robot interactions from multiple cameras, yet their outputs remain disparate video collections rather than a shared metric scene queryable across viewpoints and time. While existing 4D reconstruction methods offer a path to spatialize these predictions, independently reconstructing and merging each camera stream fails to enforce cross-view consistency. This limitation is particularly detrimental when combining moving robot-mounted cameras with fixed external views. To address this, we introduce RoGSW4RLD, a feed-forward framework that lifts synchronized multi-camera rollouts into a unified, time-queryable metric 4D Gaussian field. Rather than learning a separate geometric transition model, RoGSW4RLD directly reconstructs the visual future generated by existing world models. Its core innovation is a two-stage architecture: Stage 1 jointly forms the metric 4D field by fusing cross-view evidence with robot-specific articulated geometry and kinematics, while Stage 2 refines the field's geometry and appearance while strictly preserving the initial temporal displacements. Evaluated on 256 held-out DROID episodes, RoGSW4RLD significantly outperforms camera-wise reconstruction with calibrated merging, improving novel-view PSNR by 2.15 dB, reducing depth AbsRel by 47%, and lowering robot displacement error by 61%. These robust gains extend to action-conditioned Cosmos 3 rollouts, demonstrating that predicted video futures can be successfully translated into consistent, spatially queryable 4D metric representations.