Papers for

human-robot interaction teams

Papers whose findings have a practical use for this group, as judged from the abstract. Open a paper to read what it means in practice.

World action models improve robot and human action prediction accuracy

From World Models to World Action Models: Rethinking Next-State Prediction

Abstract: Predicting the next state is a core paradigm of World Models for modeling physical dynamics, emphasizing prediction fidelity. As World Models evolve into World-Action Models (WAMs), existing methods still fix the next state before training as RGB, a single latent feature, or a static combination of predefined targets, thereby constraining action learning to the inductive biases preserved by a particular representation. To address this limitation, we propose CF-WAM, a dynamic next-state prediction framework that samples visual, semantic, geometric, and interaction projections of the same future, standardizes them into a common video form, and supervises a unified WAM across these projections. The action-relevant constraints exposed by these projections accumulate across training steps, forcing WAM to capture the underlying state-transition structure that supports multiple projections of the same action-conditioned future. This dynamic mechanism also provides a natural cross-embodiment dynamics reference frame for Human and Robot learning. By jointly learning across different next-state parameterizations, heterogeneous Human and Robot experience can bypass appearance differences and directly contribute to shared state-transition learning, improving cross-embodiment generalization. Experiments show that CF-WAM improves both training efficiency and final control performance, while translating Human experience effectively into policy gains. CF-WAM achieves state-of-the-art performance on RoboCasa-GR1 with an average success rate of 82.50%, while reaching 82.65% on LIBERO-Plus and up to 84.00% in real-world evaluations.

Mon 28 SeptRobotics
The gist
Predicting what will happen next is important for robots and AI to understand and interact with their surroundings. The authors found that fixing how the future is predicted limits learning, so they made a method that looks at many different ways to describe the same future scene. This helps the AI learn better how actions lead to changes, even when used by both humans and robots with different appearances. Their approach improved performance in experiments, showing it helps AI control systems learn more effectively.
Open → 2609.34414v1