Humanoid robot prototype enables flexible human interaction with AI modules
Development of a Humanoid Robot Prototype for Multimodal Human-Robot Interaction
Robotics
Summary
Humanoid robots need to understand human movements, speech, and objects around them to work well with people. This paper presents a robot with moving arms and head, an expressive face, and a computer that runs several smart programs. These programs help the robot recognize gestures, identify objects, and understand voice commands. Tests show the robot can accurately move and interpret tasks over 90% of the time, making it a useful tool for robots working alongside humans. The authors suggest this robot can help others create and improve human-friendly machines.
human-robot interactiondegree of freedomgesture recognitionobject detectionspeech recognitionlarge language modelMediaPipe PoseYOLOLSTM classifierJetson module
Authors
Thang Tran Viet, Thanh Nguyen Canh, Huy Uong Gia, Phuc Dinh Van, Son Tran Duc, Ngoc Minh Do, Xiem HoangVan
Abstract
Human-robot interaction (HRI) enables intuitive and intelligent collaboration between humans and robots in real-world environments. This paper introduces a humanoid robot prototype designed as a flexible testbed for developing and integrating artificial intelligence (AI) modules in HRI tasks. The system features a 12 degree-of-freedom (DOFs) dual-arm mechanism and a 2 DOFs head with an expressive LCD screen to express facial emotions. All hardware components are controlled by a custom-designed controller board with real-time AI processing supported by an onboard Jetson module. The system incorporates three AI modules: (1) gesture recognition using MediaPipe Pose and an LSTM classifier, (2) object detection with YOLO and 3D localization, and (3) voice-command processing through speech recognition and large language model(LLM)-based semantic parsing. The platform is validated through experiments on positioning accuracy, with results showing average manipulation errors of approximately 1.83 cm. To demonstrate its versatility, experimental results show over 90% task accuracy, with gesture recognition reaching 96%, speech recognition reaching 92%. The results confirm the effectiveness of the proposed system as a reproducible and accessible humanoid platform for research and prototyping in HRI.