Semantic mapping improves urban navigation across object sizes

Multi-Scale Semantic Mapping in Urban Environments via Observation Calibration and Policy Dependence Regularization

Computer Vision and Pattern Recognition

Summary

Mapping cities for navigation is harder than indoors because objects like people and buildings vary greatly in size and distance from the observer. The authors created a new city-scale dataset to study this problem and designed a method to better handle observations of objects at different sizes and distances. They also introduced techniques to make sure the parts of their method work well together and specialize for different object categories. This helps robots or agents understand urban environments more reliably for navigation.

What this means in practice

  • For autonomous vehicle developers: Improve urban navigation by accurately mapping diverse objects from pedestrians to buildings at varying distances and sizes.$Commercial implications: Enables creation of advanced mapping systems for self-driving cars that handle complex city scenes more effectively.
  • For robotics navigation teams: Create better urban robot navigation systems by calibrating observation reliability and using specialized mapping strategies per object category.

Authors

Runling Long, Junhao Feng, Jia Wan

Abstract

Semantic mapping is fundamental to embodied navigation, yet existing methods are developed for indoor environments, where objects exhibit relatively limited scale variation and are observed from a restricted range of viewpoints. Urban environments pose substantially greater challenges: agents must map objects ranging from pedestrians to buildings while navigating large spaces with highly diverse viewing distances. These conditions introduce two key difficulties that existing datasets and methods fail to cover. First, object scale and observation distance can be severely mismatched. For example, small objects may be viewed from far away, whereas large objects may be observed at extremely close range, resulting in unreliable observation likelihoods. Second, objects with substantially different sizes and geometries require distinct mapping behaviors, which are difficult to capture with a single shared value estimator. To investigate these challenges, we introduce a large-scale urban semantic mapping dataset featuring realistic city layouts, high-fidelity rendering, and instance-level annotations spanning multiple object scales. We then propose a category-aware likelihood calibration policy that identifies and alleviates unreliable observations according to object category and viewing distance. Because the calibration and motion policies are optimized toward the same mapping objective, they may learn redundant shortcuts and become excessively coupled. We therefore introduce a mutual-information (MI) regularizer that penalizes their estimated representation dependence and encourages complementary behaviors. To better model heterogeneous mapping strategies across object scales, we further employ category-wise value estimators. We formulate their joint optimization as a Pareto optimization problem to mitigate conflicting gradients across categories.