Papers for

robot navigation teams

Papers whose findings have a practical use for this group, as judged from the abstract. Open a paper to read what it means in practice.

Gpu-accelerated collision-aware planning improves 3d gaussian splatting scenes

CollisionSplatting: Collision-Aware Motion Planning in 3DGS Scenes with Image-Conditioned Objectives and Adjustable Conservatism

Abstract: Incorporating dense visual information into motion planning remains challenging, as geometric planners rely on abstracted scene representations that discard visual richness, while learned visual models often lack geometric interpretability and computational efficiency. This paper introduces CollisionSplatting, a simple, modular, GPU-accelerated, probability-inspired distance metric with tunable conservatism that operates directly on standard 3D Gaussian Splatting (3DGS) scenes. When combined with learned image-conditioned reward functions, this metric enables joint geometric and visual planning by unifying collision-aware costs with image-space objectives. We integrate the metric into GPU-accelerated Model Predictive Path Integral (MPPI) and Rapidly-Exploring Random Tree (RRT) planners, and show on-par or better collision-classification performance compared to representative baselines while achieving substantially higher collision-checking throughput and significantly lower VRAM usage. Finally, we demonstrate the effectiveness of our metric in real-world vision-guided navigation and manipulation tasks, highlighting 3DGS as a practical bridge between rich perception and real-time motion planning.

Mon 28 SeptRobotics
The gist
Planning safe movement for robots using detailed 3D scenes is hard because traditional methods simplify visuals or are slow and hard to interpret. The authors propose a new technique called CollisionSplatting that uses a fast, adjustable way to measure distances directly on detailed 3D visual data. This method works well with AI tools that use images to guide decisions and fits nicely into popular robot path planning methods. Their approach runs fast, uses less memory, and makes smart, safe navigation possible in real-world tasks.
Open → 2609.35619v1

3D point tracking improves metric accuracy using state space models

3D Point Tracking with State Space Models

Abstract: Tracking any point of a dynamic scene in metric 3D - in absolute meters, not up to an unknown scale - underpins 3D and 4D reconstruction, robot navigation, and autonomous driving, where decisions are made in meters, not pixels. Our objective is a 3D point tracker accurate in those absolute terms and operating within a single commodity GPU, pose-free, monocular budget. Our method rests on one observation: once a point's 2D image trajectory is fixed, the quantity that governs its metric accuracy is the depth along its pixel ray. Rather than learning tracking end-to-end, we therefore compose two frozen front-ends - dense optical flow for 2D correspondence and a monocular metric-depth network for the third dimension - and learn only the residual they cannot supply: that depth, refined by a compact state space model (Mamba-3) conditioned on appearance features (DINOv3). A state space model rather than the transformers the strongest 3D trackers adopt is what makes a single-GPU budget attainable: it summarises a track in a fixed-size recurrent state whose memory cost is constant in the number of frames, whereas attention requires a key-value cache that grows linearly with them. On the TAPVid-3D minival benchmark our best configuration attains the highest absolute metric accuracy among methods evaluated under identical conditions (mean metric Average Jaccard, 0.256), exceeding strong feed-forward trackers, while a companion analysis, reproduced with each competitor's own evaluator, explains why several published trackers lose most of their accuracy under this budget.

Sun 27 SeptComputer Vision and Pattern RecognitionMachine LearningRobotics
The gist
Tracking points in 3D space accurately is important for robots and self-driving cars, which need real-world distances rather than just pixels. The authors combine two existing tools for 2D tracking and depth estimation and add a small model to improve depth accuracy over time. This approach runs on a single common GPU and performs better than other methods using similar resources. Their method helps track points precisely in meters rather than arbitrary scales.
Open → 2609.34035v1

Fast method predicts how far uncertain vehicles may move toward robots

Fast Direction-Conditioned Reachability for Motion Prediction Under Model Uncertainty

Abstract: To avoid collisions, a robot must repeatedly predict where nearby agents may move, usually with an imperfect model of their dynamics. Reachable sets provide such predictions, but computing them when the system matrices themselves are uncertain can become computationally expensive and conservative for frequent replanning. Moreover, a planner often needs to know only how far an agent can move in one particular direction, for example toward the robot, rather than the complete reachable set. We propose a direction-conditioned reachability method for linear systems with uncertain state and input matrices. Given a query direction $d$, the method selects one admissible model $(A^\star,B^\star)$ whose reachable set extends nearly as far along $d$ as the reachable set of the entire uncertain model family, and then computes the reachable set of only this model with a standard reachability solver. On an uncertain linearized bicycle model, the complete selection-and-computation pipeline is about three times faster than computing the reachable set of the full uncertain family in the CORA toolbox, while its extent along $d$ is within $5\%$ of the full family's in the reported directions. We also use the method in a closed-loop multi-vehicle simulation in which the robot queries, at each replanning step, how far each nearby vehicle can move toward it, and replans to avoid the resulting sets.

Tue 22 SeptRobotics
The gist
Robots need to guess where nearby moving things like cars will go to avoid crashing. This is hard because the robot's model of how things move isn't perfect and can be slow to calculate. The authors created a faster way to figure out how far an agent might move in just one direction, such as toward the robot, by focusing on one likely motion model instead of all possibilities. Their method is about three times faster and almost as accurate, helping robots plan safer paths around uncertain moving vehicles.
Open → 2609.27077v1

Cube-splat improves 360 degree slam tracking and mapping accuracy

Cube-Splat: High-Fidelity 360° Gaussian Splatting SLAM via Cubemap Factorization and Adjoint-Consistent Optimization

Abstract: Recent progress in 3D Gaussian Splatting (3DGS) has enabled dense visual SLAM with pinhole cameras, yet most pipelines are not designed for panoramic imagery. We present Cube-Splat, the first panoramic GS-SLAM framework that factorizes each 360° frame into a cubemap of four fixed-orientation virtual pinhole views sharing a single optical center. By designating the front face as the primary pose state, we accumulate gradients from all faces via an adjoint mapping, thereby enabling multi-face observations to coherently update a single state while strictly preserving cross-view geometric consistency. Concurrently, our mapping module densifies and optimizes anisotropic Gaussians using aggregated cubemap rays for high-fidelity, dense reconstruction. Furthermore, to rigorously evaluate panoramic SLAM under diverse and challenging conditions, we introduce SynPano, a highly scalable, photorealistic synthetic dataset featuring parameterized complex trajectories and multi-modal ground truth. Extensive evaluations on two public benchmarks (PALVIO and OmniBlender) and our SynPano dataset, collectively encompassing both indoor and outdoor scenes, demonstrate that Cube-Splat achieves state-of-the-art (SOTA) performance in tracking accuracy and reconstruction fidelity. Both the source code and the SynPano dataset are available at https://github.com/guoxf304/CubeSplat.

Fri 18 SeptComputer Vision and Pattern Recognition
The gist
360-degree cameras see everything around you, but it’s hard for computers to use these images to understand their location and surroundings accurately. The authors created Cube-Splat, a method that breaks down 360-degree images into smaller square views to better track movement and build detailed 3D maps. They also made a new dataset called SynPano that helps test these methods in different environments. Their approach works better than others on public benchmarks for both how well it tracks movement and how closely it reconstructs scenes.
Open → 2609.21347v1