GraphWrit3R generates 3D scene graphs directly from point clouds

GraphWrit3R: End-to-End 3D Scene Graph Writing

Computer Vision and Pattern Recognition

Summary

Understanding 3D environments means recognizing objects and their relationships. Previous methods required complicated steps and extra information that aren’t available in real situations. The authors created GraphWrit3R, a simpler method that takes raw 3D data and directly produces a detailed map of objects and how they relate. It works with different data types and uses a large language model to describe scenes in natural language. Their method matches or exceeds the best existing techniques without needing extra annotations during use.

What this means in practice

  • For robotics engineers: Automatically generate detailed maps of surroundings including object relationships from raw 3D sensor data for robotic navigation and interaction.
  • For augmented reality developers: Create real-time semantic scene graphs from diverse 3D inputs to enable richer and more interactive AR experiences without relying on pre-labeled objects.

Authors

Luka Milivojevic, Nikola Popovic, Sayan Deb Sarkar, Sebastian Koch, Iro Armeni, Luc Van Gool, Danda Pani Paudel

Abstract

3D scene graphs provide a structured representation of complex environments by encoding objects, their semantic attributes, and the spatial and functional relationships between them. Current approaches for 3D scene graph generation suffer from several fundamental limitations. They rely on complex multi-stage pipelines with explicit intermediate representations, making systems fragile and prone to error propagation. They assume access to ground-truth object annotations during inference, which deviates from real-world scenarios. They depend on proprietary models, hindering open-source deployment, or incur prohibitively slow inference. We present GraphWrit3R, a simple end-to-end method that takes a 3D point cloud, Gaussian Splats, or a combination of both as input, and directly outputs a complete scene graph as a structured JSON script. The graph lists all objects, their semantic attributes, and the relationships between them, while avoiding all of the above mentioned limitations. The choice of multiple input modalities is purely for versatility, allowing a single set of weights to handle diverse scenarios. Point cloud inputs are encoded via Sonata and Gaussian Splat inputs via Chorus, with both modalities projected onto a shared voxel grid and fused through a novel per-voxel contrastive alignment loss before being decoded by a large language model. As a natural consequence of the LLM, GraphWrit3R also supports open-vocabulary querying. On the 3DSSG benchmark, our method achieves state-of-the-art performance on object class, predicate, and triplet recall, outperforming methods that rely on ground-truth object annotations during inference. We further provide qualitative results and analyze different input modality configurations, contrastive loss formulations, and token fusion strategies.