MmWave radar enables full 3D body mesh without cameras

Privacy-Preserving Full-Body Meshing from mmWave Radar via Mesh Foundation Model Supervision

Artificial IntelligenceComputer Vision and Pattern Recognition

Summary

Measuring people’s body shapes and movements often uses cameras that can invade privacy. This paper shows how to use very sparse radar data instead, turning it into detailed 3D body models without any cameras involved during operation. The authors use a smart teacher-student system where a model trained on camera images teaches another model to understand radar signals. Their new approach estimates not only body poses but also how uncertain each measurement is, improving reliability while keeping people's privacy intact.

What this means in practice

  • For security system developers: Build privacy-respecting human monitoring systems that reconstruct full 3D body meshes from radar alone without cameras.$Commercial implications: Enables development of commercial secure premises monitoring solutions that respect privacy by avoiding video cameras.
  • For smart home device makers: Create home automation systems that detect detailed human body movements using existing low-cost radar sensors.

Authors

Shuxing Zhang, Yongquan Ni, Zhenyu Ding, Yawen Lin

Abstract

Millimeter-wave (mmWave) radar enables privacy-preserving human perception, but the extreme sparsity of point clouds from commercial single-chip sensors (mean ~6.5 points/frame; ~28% empty frames) has confined prior art to body-part keypoints or discrete action classification. We present a cross-modal teacher-student framework that lifts commercial radar to full-body, per-frame, metric 3D mesh reconstruction with per-joint uncertainty. Three innovations: (1) a mesh-foundation-model teacher - SAM 3D Body produces whole-body MHR ground truth (70 joints, 18,439 mesh vertices) from a single RGB frame with zero training, slashing annotation cost by orders of magnitude; (2) StudentPoseFormer - set encoding with masked attention pooling, a temporal Transformer, and a CVAE multi-hypothesis head that outputs both the pose mean and per-joint variance, honestly reporting where the radar cannot see; and (3) a multi-stage ground-truth quality pipeline (confidence gating, depth validation, temporal smoothing, bone-length consistency, bad-frame rejection) plus systematic information-lever ablations. On the public MM-Fi benchmark (same TI IWR6843 sensor, cross-subject), our full configuration reaches 7.45 cm 12-joint MPJPE, with ablations proving the causal value of point accumulation (k = 3, -0.34 cm), Doppler (-0.85 cm; -2 cm at the wrist on fast actions), and velocity loss (-0.27 cm). On our own synchronized radar + RGB-D corpus with block-level held-out splits, the pipeline achieves 21.47 cm end-to-end (per-joint hierarchy from 4.8 cm at the hip to 34.7 cm at the wrist - matching physical information limits), could be improved to 15 cm with ~30k diverse samples, and a scaling law shows sample diversity, not volume, is the binding constraint. Deployment inference is radar-only - no camera, no image.