Framework improves robot cooperation by sensing human intent accurately
A Confidence-Aware Multimodal Fusion Framework for Industrial Human-Robot Collaboration
RoboticsHuman-Computer Interaction
Summary
Industrial robots need to understand what humans want to do to help safely and efficiently. This work introduces a system that combines different signals like hand movement, gaze, and object position to predict human intentions better. The system also checks which signals are reliable in real time and adjusts how it uses them. Tests on a real robot show it works well even in tricky conditions like poor lighting or when some views are blocked. This makes human and robot teamwork smoother and more adaptable in factory settings.
What this means in practice
- •For industrial automation engineers: Enable robots to predict worker intentions reliably in assembly tasks combining multiple sensor inputs for safer, efficient human-robot teaming.
- •For robotics system integrators: Integrate confidence-aware multimodal inputs to improve robot responsiveness under varying environmental conditions like dim lighting or occlusions.
Authors
Xinyu Liu, Qiqi Dong, Boya Jia, Yi Zhang, Binbin Lian
Abstract
A confidence-aware multimodal fusion framework (CAMF) is proposed to realize reliable human intention prediction for industrial human-robot collaboration. This framework fuses four heterogeneous modalities including object 6D pose, gaze, skeletal motion and IMU-based hand motion. It embeds a confidence-trend-driven dynamic fusion mechanism into BiLSTM to adaptively balance bidirectional temporal features according to real-time modality reliability. A confidence-guided balanced learning strategy combined with a confidence freezing mechanism is further adopted to adjust network gradients dynamically, suppress noise from low-quality modalities and mitigate cross-modal learning bias. A physical platform based on the UR3 collaborative robot is built for experimental validation. Comparative results show that the proposed method reaches an intention recognition accuracy of 91.86% and outperforms existing multimodal fusion approaches in overall performance and stability. It also maintains satisfactory accuracy under low light and partial occlusion interference. In practical assembly tasks, the framework enables proactive and stable human-robot cooperation with strong environmental adaptability.