Pipeline segments and labels dual-arm robot actions with vision language
Automatic Labelling for Bimanual Mobile Manipulation
Robotics
Summary
Understanding what a robot is doing, especially when it uses two arms at once while moving, is tough. The authors created a method that breaks down a robot’s movements into meaningful parts by looking at its motion and pictures with descriptions. This method labels what each arm and the base are doing over time in a way people agree with. Their approach helps make sense of complex robot actions with words and times, making it easier to check and build on robot behavior.
What this means in practice
- •For robotics engineers: Automatically generate detailed labels for dual-arm mobile robots’ actions to improve task monitoring and policy learning.
- •For industrial automation teams: Use structured annotations from robot task data to verify and refine manufacturing tasks involving two-arm mobile manipulators.
Authors
Yupu Lu, Jia Pan
Abstract
Semantically meaningful subtask labels can provide useful contexts for long-horizon policies, but automatically identifying both reliable temporal boundaries and broad semantic descriptions for annotations remains difficult. We present an automatic labelling pipeline that assigns temporal localisation to deterministic trajectory analysis and semantic interpretation to vision-language (VL) reasoning. The pipeline segments synchronised kinematic signals into phases, performs phase-localised VL reasoning to describe the contents, and aggregates the outputs for the base, left arm, and right arm actions. We evaluate this pipeline primarily on 29 real Galaxea bimanual mobile-manipulation tasks. Repeating the VL reasoning three times first produces the same output value for 87.4% on selected tasks. A review by nine participants across all 29 tasks then judgements on the labelled phases and shows positive acceptance of temporal divisions (90.5%), body labels (90.7%), and arm labels (78.7%). The results indicate that the segmentation-VL design can produce structured annotations while preserving asynchronous bimanual behaviour, providing a basis for richer semantic subtask identification and state-based verification.