3D Human Skeleton Estimation from Single RGB Image Based on Fusion of Predicted Depths from Multiple Virtual-Viewpoints

Wen‐Nung Lie, Veasna Vann · 2023

Estimating 3D human pose from a single RGB image is still challenging in computer vision. Driven by multi-view approach, we propose a model to estimate multiple virtual-view skeletons from a single real-view image which are then fused to derive the final 3D human skeleton. Our network is composed of two stages. The first-stage is a two-streams network, where a Real-Net stream predicts 2D image coordinates and relative depth for each joint from the real-view image and a Virtual-Net stream predicts the relative depths in virtual viewpoints for the same joints. Our second-stage network consists of depth-denoising module and fusion network, where the outputs of the first-stage network are concatenated, depth-denoised, and converted into the predicted 3D human skeleton. The experimental results show that our technique has achieved a performance of MPJPE=47.48 mm, which is comparable to other state-of-the-art methods based on a much longer image sequence (9~243 frames).

Read the paper · More papers on PaperTik