ViT-Inertial Framework for Real-Time 3D Human Motion Pose Estimation
Yiliang Chen, Jie Ren, Chunlin Luo · International Journal of Image and Graphics · 2025
Real-time 3D human motion pose estimation is of great significance in fields such as virtual reality, human–computer interaction, and motion analysis. However, the complexity and variability of background environments, as well as the diversity and dynamism of human poses, lead to inaccurate estimation results. To address this, we propose ViT-Inertial framework for real-time 3D human motion pose estimation. First, the Vision Transformer is utilized to extract deep human motion features from images, and a positional encoder is used to address the issue of missing sequence position information. Then, data from visual and inertial sensors are combined through time synchronization, filtering, and transformation steps to further enhance the accuracy and robustness of human motion pose estimation. Subsequently, a human motion model is constructed, and the Kalman filtering method is used to fuse visual and inertial features, resulting in an integrated motion pose feature set. Finally, by employing a motion library and human motion prior knowledge for prediction, and combining local and global optimization methods, real-time and accurate estimation of human motion poses is achieved, enabling the tracking of continuous motion poses. Experimental results show that compared with other methods, the proposed method yields estimation results that are consistent with actual results, with the highest AUC value, joint position deviation below 5 mm, and correct key-point percentage above 0.9, demonstrating high estimation accuracy and practical application value.