Reflection aware real time slam improves indoor 3d mapping and rendering

RRG-SLAM: Real-time Reflection-aware Gaussian SLAM for Indoor Scenes

Computer Vision and Pattern RecognitionGraphics

Summary

Indoor mapping systems sometimes get confused by reflections on shiny surfaces, leading to mistakes in 3D maps or camera tracking. The authors present a system that can tell apart normal scene details from reflections by modeling them separately. This improves the quality of indoor 3D maps, the stability of camera tracking, and the rendering of new views, all while running fast enough for real-time use. Their system detects reflective planes, models reflections explicitly, and combines this information smartly during tracking and rendering.

What this means in practice

  • For roboticists: Improve indoor robot navigation and mapping accuracy in environments with reflective surfaces by separating reflections from base scene features in real time.
  • For augmented reality developers: Create more realistic indoor AR experiences by accurately handling reflections during 3D reconstruction and view rendering on mobile or wearable devices.$Commercial implications: Enables more realistic AR content that properly accounts for reflections, enhancing user immersion and appeal.

Authors

Yong Liu, Keyang Ye, Zhexi Peng, Ruixian Mei, Kun Zhou, Tianjia Shao

Abstract

We introduce the first real-time reflection-aware Gaussian SLAM system for indoor scenes. The system features a reflection-aware TSDF-Gaussian hybrid representation that explicitly separates diffuse scene appearance from reflection components. The base scene is modeled by a TSDF volume and a set of base Gaussians capturing geometry and diffuse appearance, while planar reflections are represented by reflection Gaussian groups associated with detected reflective planes. The rendering is performed in three passes: TSDF raycasting first yields surface color, depth, plane IDs and reflection masks; base Gaussians are then rendered order-independently with depth culling and combined with the TSDF output to form the base image; finally, under the guidance of the plane ID map, reflection Gaussians from different reflection groups are rasterized only into their corresponding planar regions to generate the reflection image, which is subsequently composited with the base image via the reflection mask to produce the final output. For online reconstruction, our system first estimates the camera pose through reflection-aware tracking to suppress interference of reflection-dominated regions. It then identifies reflective planes using geometric, semantic, and temporal cues, and fuses the observations into the augmented TSDF volume with reflection-aware attributes. Afterwards the base and reflection Gaussians are initialized, optimized, and pruned online to maintain both reconstruction quality and efficiency. Experiments on a variety of datasets show that our method outperforms existing SLAM systems in reconstruction quality, tracking robustness, and novel-view rendering for indoor environments with reflections, while preserving real-time performance.