Leveraging Regression and Spatial-Temporal Attention for Efficient Video-Based Human Pose Estimation
Honhlei Zu, Linhao Xu · 2024
Video-based human pose estimation is a crucial task in computer vision. Current methods typically rely on heatmaps and aggregate results from multiple frames to determine key poses. However, these methods have drawbacks, including high computational costs associated with heatmap generation and low detection efficiency, which hinder real-time video processing capabilities. In this paper, we propose an efficient method based on regression and aggregation of results from multiple frames to enhance both the accuracy and efficiency of pose estimation. Specifically, we employ a novel videomae2 backbone network to extract features across the entire video sequence, which are then fed into a multi-temporal attention network for cross-frame computation. Finally, a smoothing network is integrated for post-processing to regress multi-frame coordinates. We validate our method comprehensively with experiments on the PoseTrack datasets, demonstrating state-of-the-art results.