Feed forward monocular slam tracks motion in extreme low light
Noctif3R: Feed-Forward Monocular Real-Time SLAM for Photon-Limited Scenes on Embedded Hardware
Robotics
Summary
Robots that navigate in very dark places need to figure out where they are using just one regular camera, even when the images are almost all black and noisy. Existing methods either take too long or stop working in such dark conditions. The authors created a new method called SYS that processes low-light images quickly and reliably, tracking the robot’s movement better than previous techniques. They also made this method run efficiently on small embedded computers that robots can carry, using less energy and memory.
What this means in practice
- •For robotic system engineers: Enable robots to localize in extremely dark environments using only a single low-resolution RGB camera with low computational resources.
- •For embedded hardware developers: Implement real-time visual SLAM algorithms optimized for low-light conditions on power-efficient embedded computing platforms like the Jetson AGX Orin.
Authors
Mihir Chauhan, Aditya Uday Abhang, Kevin Biju Mathew, Aniket Bera
Abstract
Robots carrying out tasks in dark environments need to localize from a single RGB camera, in light so low that the per-pixel signal approaches the sensor's own noise, on a power-constrained onboard computer, in real time. Each of these constraints has matured pipelines, but the intersection does not. Offline low-light reconstruction now recovers structure below -4 dB but is far too slow to run in real time, while the real-time monocular systems a robot can actually carry (DROID-SLAM, DPV-SLAM, etc.) degrade or fail when SNR gets low. We measured how they fail: across the nine lowest darkness levels of our scenes, DROID-SLAM returns a full-length trajectory carrying no information about the camera's motion on all nine, VGGT-SLAM and CUT3R on eight, pi^3 on seven, and DPV-SLAM on four. We present SYS, a monocular pipeline built on a low-light feed-forward pointmap front end with an explicit match gate, which returns three tracked trajectories and no uninformative ones, at the lowest error of any method where it tracks (24-47% of the no-information ceiling against 56-73% for the strongest baseline), and at the narrowest coverage. On a real robot video take in which 86.5% of delivered frames are entirely black, every configuration of ours stops after the lit beginning, while DROID-SLAM and DPV-SLAM each emit a pose for all 1178 frames. Our method contribution is an embedded execution path for the Jetson AGX Orin: running the map, keyframes and backend at 384 pixels with tracking at 256, together with two fixes to the per-frame pose solve, is a replicated Pareto improvement, 1.28x throughput at 0.964x error on one scene and 1.42x at 0.68x on a second, with 47% less peak GPU memory and 29% less energy per pose. We evaluate on a calibrated, bit-exact regenerable noise ladder, on relabelled real-world dark exposures, and on a new dark-room video ladder recorded from a Boston Dynamics Spot robot.