Efficient 3D human pose estimation for IoT-based motion capture using Spatiotemporal Attention

Chen Zhang, LI Lu-yan, Zhihao Zhang, Yan Zhou · Alexandria Engineering Journal · 2025

With the growing demand for efficient and accurate 3D human pose estimation in fields such as virtual reality, human–computer interaction, sports analysis, and IoT-based monitoring, current Transformer-based solutions face challenges due to their quadratic computational cost as the number of joints and frames increases. To address this, we propose a 3D pose estimation network that combines Spatio-Temporal Criss-Cross Attention (STC) and a central point attention mechanism. The STC module splits the input features into spatial and temporal parts, applying self-attention to capture joint relationships within spatial frames and track dependencies across temporal frames. The central point attention mechanism uses a voxel network to refine pose regression within the central point range. By stacking multiple STC modules and introducing structure-enhanced positional embedding (SPE), our method captures spatiotemporal features and local structures. Experiments on the Human3.6M and MPI-INF-3DHP datasets show our approach achieves state-of-the-art accuracy with low computational cost, making it ideal for IoT-based monitoring and real-world applications requiring efficient pose estimation.

Read the paper · More papers on PaperTik