Visual inertial fusion improves accuracy with learned noise measurements

MAC-I$^2$: Learned Metrics-Aware Covariance for Robust Visual-Inertial Fusion in Initialization and Calibration

Robotics

Summary

Combining camera and motion sensor data helps devices know their position and movement, but often this is less reliable in tricky situations like poor lighting or moving objects. The authors created a system that learns how much to trust each type of measurement based on the current situation, instead of using fixed assumptions. Their method adjusts how it combines data from the camera and motion sensors by estimating the actual level of noise or error in the measurements. This leads to much better success in figuring out position and movement, especially during the important starting phase when setting up the system. Tests showed their approach works much better than earlier methods, especially in challenging environments.

Visual-Inertial fusionCovarianceUncertainty estimationPose estimationIMU (Inertial Measurement Unit)Feature matchingSensor calibrationState initializationRobust state estimationEuRoC dataset

Authors

Xiang Fei, Yuheng Qiu, Can Xu, Yutian Chen, Ruogu Li, Xingxing Zuo, Wenshan Wang, Sebastian Scherer

Abstract

Visual-Inertial (VI) fusion is fundamental to accurate and robust state estimation, where camera and IMU measurements are combined according to their respective uncertainties. Existing methods, however, fuse the two modalities with predefined uncertainties, regardless of how reliable each is in the local context, and thus often struggle under challenging environments involving illumination changes, dynamic objects, and textureless regions. In this paper, we present MAC-I$^2$, which achieves robust VI fusion through learned metric-aware covariance for both modalities, so that vision and IMU compete on their own merits rather than relying on predefined uncertainties. Here, metrics-aware means that each predicted covariance faithfully reflects the actual magnitude of the corresponding measurement noise. On the visual side, we propagate learned feature-matching uncertainties into pose covariances for the fusion. On the inertial side, motivated by the observation that integration error accumulates sharply at the early stage and grows slowly afterward, we design a learned IMU model with a learnable initial covariance, and propose a dedicated fine-tuning strategy on a held-out training subset to enable the metrics-aware covariance on unseen sequences. As a showcase, we build a VI initialization and calibration system, since accurate and robust initialization and calibration are the prerequisite for any reliable VI system. Experiments on EuRoC, and VBR show that MAC-I$^2$ substantially outperforms existing methods: it achieves a 99.9% initialization success rate on EuRoC, reducing gravity and velocity errors by about 60% and 42% over the strongest baseline, and maintains 80% success rate on challenging VBR sequences where baseline methods such as VINS-Mono drop below 10%.