3D Human Pose Estimation Based on Monocular RGB Images and Domain Adaptation
João Renato Ribeiro Manesco, Stefano Berretti, Aparecido Nilceu Marana · 2024
Human pose estimation in monocular images is a challenging problem in Computer Vision. Currently, while 2D poses find extensive applications, the use of 3D poses suffers from data scarcity due to the difficulty of acquisition. Therefore, fully convolutional approaches struggle due to limited 3D pose labels, prompting a two-step strategy leveraging 2D pose estimators, which does not generalize well to unseen poses, requiring the use of domain adaptation techniques. In this work, we introduce a novel Domain Unified Approach called DUA, which, through a unique combination of three modules on top of the pose estimator (pose converter, uncertainty estimator, and domain classifier), can improve the accuracy of 3D poses estimated from 2D poses. In the experiments carried out with SURREAL and Human3.6M datasets, our method reduced the mean per-joint position error (MPJPE) by 44.1 mm in the synthetic-to-real scenario, a quite significant result. Furthermore, our method outperformed all state-of-the-art methods in the real-to-synthetic scenario.