System preserves individual views for better 3d object searches
EviSplat: Preserving Multi-View Evidence in 3D Gaussian Splatting for Open-Vocabulary Segmentation
Computer Vision and Pattern RecognitionArtificial Intelligence
Summary
Figuring out where objects are in 3D scenes from text questions can be tricky because the same object looks different from various angles. Many current methods mix pictures of an object from all views before knowing what will be asked, which can cause important clues to be lost. The authors introduce a system called EviSplat that keeps separate clues from each view instead of mixing them early. When a question arrives, EviSplat picks the best clues for that specific question, improving how well it finds and segments objects. Their tests show this method works better than others because it keeps flexible evidence until needed.
What this means in practice
- •For augmented reality developers: Enable AR apps to accurately find and highlight objects from text descriptions by retaining multiple views until search time.
- •For robotics perception teams: Improve robot scene understanding by preserving and selectively using visual clues from different viewpoints for flexible object queries.
Authors
Sungho Moon, Kota Shimomura, Junwoo Park, Wonhyeok Choi, Seunghun Lee, Takayoshi Yamashita, Sunghoon Im
Abstract
Open-vocabulary 3D scene understanding enables object localization and segmentation from free-form text queries without a fixed category vocabulary. Many recent methods build on 3D Gaussian Splatting and consolidate multi-view observations, such as masked crops from individual views, into language features or compact object descriptors before the query is known. However, observations of the same object vary across viewpoints and are not equally informative: some reveal cues relevant to a particular query, whereas others provide incomplete or misleading evidence. Pre-query consolidation can therefore suppress cues on which a later query depends. We introduce EviSplat, which preserves individual observation features as evidence for later text queries. EviSplat retains individual observation features within class-agnostic 3D instances that represent objects, object parts, or background regions. It also learns, for each Gaussian, a distribution describing which visual appearances its observations support. Given a text query, EviSplat scores each instance using its most relevant observations. It then computes a score for each Gaussian by combining instance-level relevance with locally supported evidence, weighted by how often and how unambiguously that Gaussian was observed. Different queries can thus draw on different visual cues from the same preserved evidence. Experiments across diverse datasets and evaluation protocols demonstrate state-of-the-art performance, supporting the benefit of preserving multi-view evidence until query time and aggregating it according to the query.