Learning in-the-wild Temporal 3D Pose Estimation from MoCap Data

Doruk Cetin · Repository for Publications and Research Data (ETH Zurich) · 2020

Recovering 3D poses from 2D representations has been a challenging task due to its inherent ambiguity and complexities of the human motion.Lack of annotated data makes it even harder to train neural networks for 3D pose estimation task.In this work, we propose a temporal two-stage architecture to estimate sequences of 3D poses from 2D joint detections, for which we use only 3D motion capture data without paired images for training.More specifically, we generate paired examples by projecting augmented 3D poses to 2D, on-the-fly.We modify our inputs during training through noise and masking to obtain models robust to 2D detection errors.Our approaches utilize the relative simplicity of augmenting 2D and 3D poses rather than images that lie in higher dimensions.Resulting framework employs a simple model, trained on poses without paired images, that can achieve competitive performance on common evaluation scenarios.We present our work as a baseline for the task at hand, discussing further directions for building upon the proposed methodology.

Read the paper · More papers on PaperTik