Automatic 3D Hand Pose Estimation Based on YOLOv7 and HandFoldingNet from Egocentric Videos
Van-Hung Le · 2022 RIVF International Conference on Computing and Communication Technologies (RIVF) · 2022
3D hand pose estimation problem still contains many challenges such as high degree-of-freedom (high-DOF) of 3D point cloud data, the obscured data, the loss of depth image data, especially the data obtained from the first-person viewpoint. In this paper, we propose an automated method for 3D hand pose estimation on detectable hand point cloud data collected from Egocentric vision. Our approach is the combination of YOLOv7 and HoldFoldingNet which are the current highest performing CNNs for hand detection and 3D hand pose estimation problem. We evaluate two stages in the proposed method on the FPHAB dataset, (1) the first stage is the evaluation of hand detection based on the pre-trained model of YOLOv7 and its variants, and (2) the second stage is to evaluate the result of 3D hand estimation on the HoldFoldingNet with the annotation box and the detected box and compare them with the state-of-the-art methods. The result of detecting hand action is the lowest (P=97% - thresIOU= 0.95). The 3D hand pose estimation results using HoldFoldingNet on the FPHAB dataset have the lowest error of 19.98mm, 20.4mm, respectively when testing on the annotation box and the detected box of the 5th configuration,