LRC-JEPA improves planning by separating dynamics from context in world models
LRC-JEPA: Disentangling Dynamics and Residual Context for Efficient World Models
Machine LearningArtificial IntelligenceRobotics
Summary
Planning systems that predict future outcomes often struggle because they must remember both how things change and what things look like. The authors present LRC-JEPA, a new model that smartly splits these two needs into separate parts: one focusing on how the world moves, and the other on stable background details. This makes planning more accurate and efficient, especially in complicated scenes. Tests in games and real-world data show LRC-JEPA delivers better results with fewer resources than previous models.
What this means in practice
- •For robot control teams: Use LRC-JEPA to improve robots’ ability to plan actions by separating changeable states from stable background information.
- •For autonomous vehicle developers: Incorporate LRC-JEPA for faster and more reliable planning in complex driving environments by efficiently handling scene dynamics and context separately.
Authors
Luzhe Huang, Lei Chu, Jingyi Liang, Yuhuan Zhao
Abstract
Compact JEPA world models enable efficient latent-space planning, but low-dimensional representation trained under reward-free self-supervision must encode both action-conditioned dynamics and predictable visual context. This competition can entangle controllable state with high-rank nuisance appearance and degrade planning as scenes become more complex. We introduce LRC-JEPA, a lightweight end-to-end world model that routes information into a compact predictive latent $\mathbf{z}$ and learned-query residual-context embeddings $\mathbf{u}$. Only $\mathbf{z}$ is propagated by the dynamics model and used for planning, while $\mathbf{u}$ captures temporally persistent information for cross-attention reconstruction; a differentiable residual connection encourages the latent to retain complementary dynamic content. Under explicit assumptions, we show that the resulting representation is sufficient, minimal, nuisance-invariant, and disentangled. Across four simulated control environments, LRC-JEPA improves average planning success over a parameter-matched JEPA baseline by 9 percentage points and matches or exceeds substantially larger pretrained models. On the real-world Bridge-v2 set, its 5.5M-parameter active encoder outperforms DINO-WM (22.1M) and V-JEPA2 (303.9M) encoders while also enabling faster planning. Physical-state probes, reconstruction interventions, and ablations confirm the effectiveness of LRC-JEPA's representation disentanglement.