Enhancing Motion Reconstruction From Sparse Tracking Inputs With Kinematic Constraints

Xiaokun Dai, Xinkang Zhang, Shiman Li, Xinrong Chen · IEEE Transactions on Automation Science and Engineering · 2024

In virtual reality, there is a growing demand for reconstructing accurate full-body 3D avatars from the sparse motion captured through head-mounted displays and hand-held controllers. However, due to the limited information from sparse inputs, precisely reconstructing full body poses is an ill-posed and challenging task. Existing methods often exhibit notable errors in lower body poses, which results in unrealistic poses and occasional floor penetration artifacts. To address the above issue, a MLP-based model with Kinematic Constraints and Temporal Diversity (KCTD) was proposed for full body poses reconstruction, which incorporates Kinematic Constraints Hierarchical Decoder with Temporal Diversity Awareness Module and a Generative Feedback Module to further improve the accuracy of the reconstruction of the full body poses. Specifically, the potential constraints of human kinematic chain are incorporated into the model through a hierarchical decoder, which elevates overall precision through the interaction of the human kinematic chain. Then, a temporal diversity awareness module is integrated into the hierarchical decoder to help the model capture information at different frequency in the time domain. In addition, a generative feedback module is imposed on leg poses reconstruction to further improve its accuracy without increasing the model’s inference time. Test results on the AMASS dataset demonstrate that, the proposed model effectively improves the reconstruction accuracy of the full-body poses with the mean per joint rotation error and position error of of 2.60 and 3.62 respectively, which surpasses the state-of-the-art methods. Particularly, the proposed model can alleviate irrational poses in the lower body and reduce the floor penetration artifacts. Note to Practitioners—The motivation behind this manuscript stems from the necessity for precise full pose reconstruction in certain virtual reality applications. While existing methods have significantly contributed to the reconstruction of full-body poses based on sparse inputs, they have not taken tailored measures to improve accuracy in the lower body. Consequently, this paper introduces a full human body pose reconstruction model grounded in constraints from the human kinematic chain, featuring a temporal diversity awareness module. This model excels in reconstructing more accurate complete human body poses using sparse inputs from the head and hands, notably addressing issues such as irrational lower-body poses and floor-penetrating artifacts. Benefiting from state-of-the-art precision, the proposed approach has the potential to significantly elevate user experience and immersion in virtual reality applications, particularly in scenarios such as automated task training and simulation of hazardous conditions.

Read the paper · More papers on PaperTik