World models adapt safety checks to reduce robot errors
When World Models Lie: Adaptive Safety Analysis Under Wrong Imaginations
RoboticsArtificial Intelligence
Summary
Robots use world models to predict what will happen next, helping keep them safe. But when these models are wrong, the safety checks might also be wrong, leading to risky decisions. The authors introduce a method that watches for when the model's guesses don't match real observations and adjusts safety warnings accordingly. This approach keeps robots cautious when needed but not overly careful when the world model is accurate.
What this means in practice
- •For robotics engineers: Improve robot controllers to better handle unexpected environment changes by adapting safety decisions based on direct model errors.
- •For autonomous vehicle developers: Enhance safety systems in self-driving cars by dynamically adjusting safety boundaries when the vehicle's internal models mispredict the environment.
Authors
John Cao, Somil Bansal
Abstract
World models offer a powerful substrate for safety reasoning in high-dimensional robotic systems, but they are also fallible: their predictions can be biased, miscalibrated, or confidently wrong. This creates a central challenge for latent-space safety filters, which often learn Hamilton-Jacobi safety value functions on the dynamics of a world model. If the world model is incorrect, the resulting value function can inherit its errors and produce overconfident safety estimates. Existing latent safety filters often rely on auxiliary signals such as ensemble disagreement or value-target consistency residuals for adaptation, but these signals can remain small even when the world model's predictions deviate from observations. We propose an adaptive latent safety filter that calibrates safety reasoning using directly observed world-model error. Our method uses Adaptive Conformal Inference to construct online uncertainty sets from discrepancies between predicted and observation-inferred latent states, then evaluates safety pessimistically by minimizing the learned value function over these sets. This allows the filter to remain minimally conservative when the world model is accurate, while becoming more cautious when observations reveal model mismatch. We provide a finite-time coverage guarantee for the adaptive uncertainty radius. Through simulation and hardware experiments, we show that our method significantly reduces failures relative to state-of-the-art latent safety filters while preserving task completion.