Training method improves long-term context understanding in world models

Shaping Persistent Representations from Independent Interactions

Machine Learning

Summary

When computer models try to learn how things change over time, they usually focus on what's happening right now, missing longer-lasting features that stay the same across different experiences. The authors created SPRII, a new way to train these models so that they can better gather and use persistent information by comparing different experiences that share hidden common traits. Their tests on many different tasks, from physics simulations to robotics, showed that SPRII helps models remember important long-term information and perform better at predicting and understanding tasks. This approach doesn’t need explicit labels about the shared features, instead it uses relationships between experiences to guide learning.

What this means in practice

  • For robotics engineers: Improve robot learning by enabling models to retain and use long-term information from multiple task runs without needing explicit property labels.
  • For machine learning developers: Enhance predictive models in physical simulation and control by organizing persistent features across interaction data to boost accuracy.

Authors

Ji Dai, Quan Fang, Junyu Gao, Rongfeng Guo, Haoyan Rong, YipingHuang, Yongxi Li

Abstract

World models learn environment dynamics from interaction experience. These dynamics depend on the current state and actions, as well as on properties that persist across interactions. Yet standard predictive training can reduce error using local evidence alone, without organizing persistent information into reusable context. We introduce SPRII, a training principle that uses relations between interactions as weak supervision for persistent context while retaining the learner's native objective. For example, different trajectories of the same system share persistent properties even when their states and actions differ. SPRII uses such relations to guide context learning without numerical property labels. Two composable components encourage contexts from related interactions to agree (Align) and use one interaction's context to predict another's future (Cross). Our analysis distinguishes three linked questions: what persistent information is accessible in the learned context (Formation), how that context influences a fixed predictor (Use), and whether it reduces task error (Value). Success at one stage does not guarantee success at the next. Controlled experiments show that more reliable relations improve representation organization, but adding a shared-property constraint can reduce access to a property that remains shared. Context substitutions change predictions at fixed model weights, while the benefit from history depends on prediction horizon and readout. Evaluations span thirteen settings, including controlled physical systems, public dynamics tasks, robotic and tactile data, and partner interaction, across multiple learner families. Relative to the corresponding baselines, SPRII yields average gains of over 10% in downstream task performance and over 15% in persistent-property readout. The project page is available at https://persistent-learning-review.netlify.app/interactive.html.