Video-Based 3D Human Pose Estimation Research
Siting Tao, Zhi Zhang · 2022 IEEE 17th Conference on Industrial Electronics and Applications (ICIEA) · 2022
At present, most of the video-based 3D human pose estimation methods process the video frame by frame, ignoring the timing information of the upper and lower frames of the video. When there is a single-frame estimation error caused by self-occlusion, the occlusion part will be Some strange movements appear due to information loss, and the estimated results make it difficult to produce accurate motion sequences. Therefore, considering the continuity of the upper and lower frames of the video, this paper adopts the HMR algorithm and introduces a bidirectional long-term memory network (BI-LSTM) in the encoding part to process the video sequence, and then outputs the parameters of the human body model through the SMPL parameters and inputs them into the discriminator for confrontation, train. In order for the discriminator to also take into account the semantic information of the temporal context, the discriminator in this paper also introduces a BI-LSTM network to improve the authenticity of the generated poses. Experimental results show that the accuracy of the model is improved when self-occlusion occurs.