FlexTrack3D: Advanced Single-Camera 3D Human Pose Tracking With FlexPoseNet and ZoeDepth
Yingying Chen, Zhitao Li · IEEE Access · 2024
At present, the research of three-dimensional human pose tracking mainly focuses on the multi-camera system, but lacks a tracking algorithm for monocular cameras. Therefore, a 3D human pose tracking algorithm, FlexTrack3D, is proposed, which can track human 3D pose on monocular video. FlexTrack3D innovatively combines 2D human pose detection and monocular depth estimation technology. By integrating pixel coordinates of human keypoints provided by FlexPoseNet and depth data generated by ZoeDepth algorithm, FlexTrack3D can accurately model keypoints in three dimensions and accurately track their dynamic trajectories. FlexTrack3D shows the pose and trajectory of human motion data collected in indoor and outdoor environment through 3D point cloud model, which proves its accuracy and robustness. For tracking keypoints of human, a lightweight human pose detection model, FlexPoseNet, and a parameter sharing detector with multi-scale feature information are proposed, which strengthens the learning of key features on the premise of reducing the parameters of baseline model YOLO-Pose. Multi-Head Self-Attention module and Deformable Convolution Networks module are introduced into the backbone of FlexPoseNet model, which enhances the modeling ability of the network to the target’s receptive field and improves the modeling flexibility of the network to the target structure. SlimNeck structure is introduced into the neck of the model, which effectively reduces the calculation and parameters of the model while maintaining the detection accuracy. In the experiment of COCO-Pose data set, the processing speed of FlexPoseNet is increased by 42.9% per second compared with the baseline model YOLO-Pose, and the detection accuracy of mAP50 and mAP50-95 is increased by 12.46% and 10.21%, which further improves the accuracy of the model on the premise of maintaining performance.