Selective sampling and temporal positional encoding for monocular video-based 3D human pose and shape estimation
Yupeng Hou, Han Wen, Guangping Zeng · 2025
This study introduces a novel framework for video-based 3D human pose and shape estimation, termed Selective sampling and Temporal Positional Encoding (STPE). Our method leverages selective sampling and advanced positional encoding to tackle the temporal complexities of video data and the high cost and scarcity of annotated datasets. Inspired by the Masked Autoencoder (MAE), our approach adopts a selective sampling strategy that efficiently captures the essential dynamics of human motion from partial views, significantly reducing reliance on continuous frames. The framework incorporates Rotary Position Embedding (RoPE), using rotational angles to simplify positional encoding. This innovation decreases model complexity and boosts learning effectiveness. We also introduce randomized index positions during training, introducing variability and enhancing generalization across various datasets and motion patterns. Our model, validated on standard datasets like 3DPW, MPI-INF-3DHP, and Human3.6M, shows enhanced performance in accurate and robust 3D pose and shape capture compared to existing methods. Our results demonstrate that strategic frame sampling and sophisticated positional encoding can significantly improve accuracy and robustness of video-based pose estimation systems.