GLAM model improves robot exploration and navigation in virtual spaces
GLAM: Training a latent world model over global spatiotemporal memory for active exploration and navigation
Robotics
Summary
Robots need to remember what they've seen and predict what might happen next to explore and navigate effectively. The authors developed GLAM, a model that helps robots create and use a map-like memory to plan where to go next. This system can predict future maps and navigation steps jointly, so the robot can better understand both its surroundings and its goals. Tests in virtual environments showed that GLAM NAV, the full navigation system built with GLAM, was more successful at reaching targets than some previous methods.
What this means in practice
- •For robotics engineers: Create robots that can better explore and navigate complex indoor environments by predicting future maps and navigation points jointly.
- •For virtual environment developers: Build more effective navigation agents in simulated scenes using global spatiotemporal memory for planning.
Authors
I-Tak Ieong, Ruizhi Feng, Zhaoyang Lu, Yifei Cao, Jiayao Zhao, Leon Li, Senhua Zhu, Wenbo Ding
Abstract
Active exploration and semantic navigation require an embodied agent to build memory from partial observations, predict how the evolution of observed spatial memory may support future motion, and convert that prediction into actionable plans. We present GLAM, a goal-conditioned latent world model trained over global spatiotemporal memory, and GLAM NAV, the complete navigation system built around it. Given historical map tokens, a navigation goal, and the current robot pose, GLAM jointly predicts future map representations and robot-centric waypoint latents, allowing future spatial context and navigation intent to be inferred in a shared representation space. The model follows a JEPA-like latent prediction paradigm, operates directly on map-level latent tokens rather than RGB reconstruction, and uses a pretrained waypoint encoder-decoder to supervise and decode navigation plans within GLAM NAV. Training data are collected by replaying ObjectNav expert trajectories in Habitat over HM3D v0.2 scene assets and slicing them into multi-timescale prediction samples. On a controlled HM3D-ObjectNav subset reproduction setting, GLAM NAV improves over a reproduced BSC-Nav baseline in both success rate and success weighted by path length.