A Diffusion-Based Framework for 3D Human Pose and Shape Estimation from Monocular Videos
Ni Zhao, Na Lv · 2024
Estimating 3D human pose and shape from monocular videos is a challenging task due to inherent ambiguity and occlusion, which often lead to inaccurate predictions with high uncertainty. Despite significant progress in 3D pose and shape estimation from a single RGB image, achieving accurate and temporally coherent human mesh sequences from monocular videos remains a challenging endeavour. Inspired by the recent success of diffusion models in generating high-quality outputs with low uncertainty through progressively denoising noisy inputs, we propose a novel diffusion-based framework for 3D human pose and shape estimation. This framework formulates the human mesh recovery task as a reverse diffusion process. During training, it diffuses SMPL pose parameters from ground-truth distributions into input-specific distributions and learns to reverse this process. By leveraging the capacity of diffusion models to reduce noise progressively, our method effectively addresses the inherent ambiguity of this monocular task, producing accurate and smooth human mesh sequences from videos. Comprehensive experiments demonstrate that the proposed method significantly outperforms previous video-based methods in both per-frame 3D pose and shape accuracy and temporal coherence on widely used benchmarks, including 3DPW and Human3.6M datasets.