A Conditional Diffusion Model for 3D Human Pose Estimation

Zhangmeng Chen, J. P. Dai, Junjun Pan · 2024

Monocular 3D human pose estimation is a highly challenging task due to inherent depth ambiguity and occlusion. In response to these challenges, many previous works have utilized temporal information to enhance the richness of information. However, such methods can only derive a specific 3D pose from a single image, which fails to address ambiguity. Fortunately, denoising diffusion probabilistic models have recently achieved success in various computer vision tasks (e.g., generating images and image segmentation). Diffusion-based methods produce a set of possible result for each image. Inspired by this, we explore a novel Conditional Diffusion-based 3D human Pose estimation framework (CDiffPose) that formulates 3D human pose estimation as a 2D keypoints-guided 3D pose synthesis problem. We utilize a pre-trained model to represent 2D keypoints information as conditioning information. Subsequently, we inject the conditional information into the denoising model. Our denoising model (Denoiser) adopts a cascaded approach involving GCN and Transformers. Finally, the model generates multiple 3D candidates poses for a single 2D keypoint to alleviate the inherent depth ambiguity. Extensive experiments are conducted on Human3.6M benchmark, our proposed method demonstrates competitive performance.

Read the paper · More papers on PaperTik