Millimeter-Wave Radar Human Pose Recognition Based on Multi-Scale Feature Fusion Networks

Human posture estimation (HPE) technology has garnered increasing attention from domestic researchers due to its extensive application scenarios and research value. Compared to HPE techniques based on visual sensors and wearable sensors, millimetre-wave radar-based technology offers distinct advantages including non-contact operation, device-free functionality, privacy protection, and strong penetration capabilities, rendering it a focal point for researchers. However, current research on human pose recognition using millimeter-wave radar still faces challenges such as suboptimal recognition performance and system implementation. To enhance the system's recognition capabilities, this paper proposes a cross-modal training framework that leverages high-precision 2D key points from synchronized RGB imagery as pseudo-ground-truth to supervise radar-only inference. To mitigate domain gaps between camera and radar modalities (e.g., perspective distortion, scale inconsistency, and occlusion variance), we introduce a dual-branch preprocessing pipeline with modality-specific normalization and a learnable cross-modal alignment module based on cycle-consistent projection. This framework incorporates a Multi-Scale Convolutional Fusion Network (MSCFN) to fuse multi-scale radar features from horizontal and vertical sensors, and a Structured Prediction Optimizer (SPO) based on Conditional Random Fields (CRF) to improve the confidence heatmap of predicted key points. Experiments on HuPR show AP=63.2, AP50=97.1, AP75=73.8, PCK@0.05 = 83.7, MPJPE=81.4mm, outperforming RF-Pose by +21.6 AP and prior radar SOTA by +4.3 AP.

Paper

Full text

PDF

Millimeter-Wave Radar Human Pose Recognition Based on Multi-Scale Feature Fusion Networks

Semantic Scholar · 2025

Abstract

Human posture estimation (HPE) technology has garnered increasing attention from domestic researchers due to its extensive application scenarios and research value. Compared to HPE techniques based on visual sensors and wearable sensors, millimetre-wave radar-based technology offers distinct advantages including non-contact operation, device-free functionality, privacy protection, and strong penetration capabilities, rendering it a focal point for researchers. However, current research on human pose recognition using millimeter-wave radar still faces challenges such as suboptimal recognition performance and system implementation. To enhance the system's recognition capabilities, this paper proposes a cross-modal training framework that leverages high-precision 2D key points from synchronized RGB imagery as pseudo-ground-truth to supervise radar-only inference. To mitigate domain gaps between camera and radar modalities (e.g., perspective distortion, scale inconsistency, and occlusion variance), we introduce a dual-branch preprocessing pipeline with modality-specific normalization and a learnable cross-modal alignment module based on cycle-consistent projection. This framework incorporates a Multi-Scale Convolutional Fusion Network (MSCFN) to fuse multi-scale radar features from horizontal and vertical sensors, and a Structured Prediction Optimizer (SPO) based on Conditional Random Fields (CRF) to improve the confidence heatmap of predicted key points. Experiments on HuPR show AP=63.2, AP50=97.1, AP75=73.8, PCK@0.05 = 83.7, MPJPE=81.4mm, outperforming RF-Pose by +21.6 AP and prior radar SOTA by +4.3 AP.

Similar papers

© 2026 NYSGPT2525 LLC