Privacy-preserving human robot interaction using millimeter wave radar

mmHRI: Towards Privacy-Preserving Human-Robot Interaction with Millimeter-Wave Radar

RoboticsComputer Vision and Pattern Recognition

Summary

Robots that help people often use cameras to understand gestures and commands, but cameras can invade privacy in places like hospitals. The authors created a system that uses millimeter-wave radar, which can sense motion without capturing detailed images, protecting privacy. Their system combines different radar data types and remembers past signals to better understand human actions and poses. This lets robots deliver and retrieve objects even when they can't see the person clearly, such as behind a curtain.

What this means in practice

  • For hospital robotics teams: Operate assistive robots in sensitive environments without using cameras, ensuring patient privacy while enabling gesture-based commands.
  • For restaurant automation providers: Develop service robots that can interpret non-verbal human commands in visually occluded areas without video monitoring.

Authors

Junqiao Fan, Yuxuan Hu, Bofan Lyu, Yanshuo Lu, Pengfei Liu, Jiarui Zhang, Fangqiang Ding, Lihua Xie, Gen Li, Jianfei Yang

Abstract

Assistive robots increasingly operate in many human-centered environments and perform various human-robot interaction (HRI) tasks, such as object delivery. However, most existing HRI systems rely on RGB cameras that continuously observe humans to respond to non-verbal commands, such as hand gestures. This raises privacy concerns in privacy- critical environments, such as hospital wards or restaurants, where direct camera observation of humans is restricted. To develop privacy-preserving HRI, we leverage millimeter-wave (mmWave) radar, which can sense human motion through privacy barriers without identifiable imagery. We propose mmHRI, the first multi-modal robot manipulation framework that achieves mmWave radar-guided privacy-preserving HRI. mmHRI introduces two key designs to mitigate the sparsity and temporal inconsistency of radar data in cluttered robot manipulation environments. First, we propose a dual-stream architecture that jointly learns from unfiltered raw radar tensors and radar point clouds to estimate both human actions and 3D poses. To mitigate signal inconsistency, mmHRI further incorporates a memory-based state-space model (MSSM) that retains historical radar features to reduce abrupt changes in pose/action. These estimated human states are then converted into structured textual robot instructions, which control a vision-language-action (VLA) policy for closed-loop robot manipulation and human-aware reactions. Our evaluation covers human action recognition and closed-loop delivery and retrieval. In the privacy-preserving curtain setting, mmHRI achieves 85.09% action-recognition accuracy, outperforming existing radar-based alternatives. Robot trials further demonstrate successful delivery and retrieval under visual occlusion, with stable task performance across unseen subjects, clutter configurations, and environments.