Mobile robot manipulation improved by seeing coordinating and imagining

MM-ABC: Towards Generalist Mobile Manipulation via Seeing, Coordinating and Imagining

Robotics

Summary

Mobile manipulation lets robots move around while using their arms to do tasks, but it’s hard because the robot must understand its surroundings as it moves and plan arm and base actions together. The authors created MM-ABC, a system that helps robots see their environment better, coordinate arm and base movements, and imagine future states to improve planning. They trained MM-ABC on a huge variety of robot data and tested it on multiple benchmarks and real robots, showing better success in completing tasks than previous methods.

What this means in practice

  • For robotics engineers: Build mobile robots that better coordinate movement and manipulation for household and industrial tasks using MM-ABC’s integrated perception and control model.
  • For automation system designers: Design automated mobile manipulators capable of improved spatial reasoning and multi-step task execution by incorporating MM-ABC’s coordinated arm-base collaboration approach.
  • For assistive robot developers: Develop assistive robots that navigate and manipulate objects more reliably in home or care settings using MM-ABC’s future state imagination and coordination techniques.$Commercial implications: Enables assistive robot products that perform complex, coordinated tasks in dynamic environments, improving reliability and autonomy.

Authors

Qiwei Liang, Guangyu Chen, Shaolong Zhu, Zikuan Xiao, Jinxuan Lu, Yifan Xie, Renjing Xu, Wenbo Ding, Tianxing Chen

Abstract

Mobile manipulation extends robot interaction beyond a fixed kinematic workspace by making the reachable region itself controllable. This flexibility introduces two central challenges: spatially grounded perception under continuous ego-motion and coordinated control of heterogeneous arm and base actions. Existing approaches strengthen geometry through explicit 3D representations or predictive world models, and often decouple mobility and manipulation into separate action streams. We argue that effective mobile manipulation requires not only decoupling, but also representations that support efficient cross-stream collaboration. We present MM-ABC, a foundation model built around Seeing, Coordinating, and Imagining Arm-Base Collaboration. MM-ABC combines sparse multi-level VLM features for spatial perception; a training-only future branch that uses world imagination and geometric intent as extra supervision, strengthening perception and manipulation-intent prediction and improving the overall learning signal; and MM-APT, which coordinates separate manipulation and mobility streams through masked joint attention and clean-action x-prediction. In controlled ablations, replacing clean-action prediction with velocity prediction lowers success on RoboCasa365 composite-seen tasks from 32.8% to 29.2%, and removing future supervision or multilevel conditioning causes larger drops. We pretrain MM-ABC on 5,000+ hours of heterogeneous robot data spanning 400K+ episodes, 12 datasets, and 17 embodiments. Experiments cover EBench, RoboCasa365, ManiSkill-HAB, LIBERO, LIBERO-Plus, and real-world mobile manipulation. MM-ABC achieves 44.71% success on EBench, 61.2% on RoboCasa365, 99.1% on LIBERO, 82.8% on LIBERO-Plus without perturbation training, and 83% mean success on five real-world tasks.