Methods for Estimating 3D Human Pose and Location

Márton Véges · Eötvös Loránd Tudományegyetem · 2022

Methods for Estimating 3D Human Pose and Location The goal of 3D human pose estimation is to predict the location of the joints of a person, based on an image or video input. The joint locations are expressed in a camera centric coordinate-system. The task has several potential applications, including motion capture, behavior and sports analytics. With the aid of neural networks, huge leaps were made in recent years in terms of performance, reaching very low errors on studio datasets. However, on in-the-wild datasets, the error rate still lags behind, being often twice as large. From a practical point of view, it indicates that more work is needed before these methods can be used in the real world. The underlying reason is simple: it is quite hard to create a 3D pose annotated dataset as an expensive multi-camera system with time-of-flight sensors is needed for accurate measurement. As a result, many datasets have a constrained setting with limited diversity in background, camera views and performed actions. Considering the data-hungriness of deep learning methods, alternative architectures must be found that do not need large amount of training examples. Another problem of current pose estimation methods is their restricted scope: they predict the pose relative to the hip only, ignoring the localization of the person. When only a single person is present in the image that is not an issue, but with multiple people that information might be needed. In this thesis, I propose four methods to solve these problems. The first method tackles the issue of few cameras in the training set that leads to overfitting. A siamese architecture is introduced that learns an equivariant embedding with regards to camera views. The equivariant property ensures that the network has good results on unseen camera poses, even without augmentations. The second method solves the problem of naive localization used in previous works. Those approaches employ a Perspective -n -Point (PnP) method that needs accurate 2D and 3D poses for localization. When either of those is erroneous, the PnP algorithm diverges resulting in an incorrect placement of the pose. By using a direct estimation of the location, the results become more stable. Next, the inclusion of RGB-D datasets is proposed as an auxiliary training database for pose estimation. The depth maps accompanying the images provide a weak training signal. The method shows improved prediction performance, especially for the localization task. Finally, an approach is introduced to solve temporary occlusions in videos. Temporal methods return an incorrect pose when a person is occluded, even if it is for a short time, when the neighboring frames hold enough information to interpolate the pose inbetween. The proposed algorithm can be applied as a refining step after any temporal method, correcting its predictions. I prove the accuracy of the algorithms with extensive quantitative tests.

Read the paper · More papers on PaperTik