MDST: 2-D Human Pose Estimation for SISO UWB Radar Based on Micro-Doppler Signature via Cascade and Parallel Swin Transformer
Xiaolong Zhou, Tian Jin, Yongpeng Dai, Yongping Song, Kemeng Li, Shaoqiu Song · IEEE Sensors Journal · 2024
This paper introduces the Human Pose Estimation based on the Single-Input Single-Output (SISO) Ultra-Wideband (UWB) Radar (HPSUR) benchmark, a pioneering approach in human pose estimation integrating motion capture technology based on SISO UWB radar sensors. The HPSUR dataset, consisting of 311,963 data frames, was meticulously assembled using cross-calibrated SISO UWB radar sensors in conjunction with the Noitom Perception Neuron 3 (N3), specifically designed for radar-based human pose estimation. This dataset captures diverse movements from five subjects of varying physical characteristics, performing four distinct categories of actions. In addition to establishing this comprehensive benchmark, our research proposes an innovative framework for 2D human pose estimation based on SISO UWB radar. The framework leverages the processing of Micro Doppler (MD) signatures through a unique combination of cascade and parallel Swin Transformers. The MD signatures, reflective of human kinematics, form the basis for a novel methodology in posture identification, thus enhancing the perception of human postures. Addressing the challenge of managing long-range dependencies due to the high sampling rates of radar devices, we introduce the Micro Doppler Swin Transformer (MDST) network. This novel transformer incorporates Window-based Multi-head Self-Attention (W-MSA) and Shifted Window-based Multi-head Self-Attention (SW-MSA) models to capture the inner-frame and intra-frame aspects of the MD signature adeptly. Furthermore, the study integrates an Inverted Feature Pyramid Network (IFPN) for an efficient multi-scale feature representation, enriching the feature pyramid with high-level semantics. Our extensive experimental analysis, conducted on the HPSUR benchmark, demonstrates the significant enhancement in the accuracy of human pose estimation offered by the proposed MDST network. This improvement is consistently observed across six MDST variants under various conditions involving diverse subjects and postures, showcasing the robustness and generalizability of our approach.