3D hand pose estimation using skeleton depth map as intermediate supervision

Rina Wu, Tianqiang Zhu, Yi Sun · 2025

3D hand pose estimation from a single depth image is a challenging problem for human-computer interaction. In order to balance the efficiency and the accuracy in modeling the high nonlinear mapping from a 2D depth image to 3D joint coordinates, two strategies are proposed. One is a new data pre-processing method, which makes the distribution of training set and test set in 3D space more concentrated and closer. The other is using the skeleton depth map as a bridge to connect each finger and its joints. Both of the above strategies can effectively simplify the task difficulty and help to build a lightweight network. For each skeleton depth map, we insert a stable triangle in each skeleton of a finger to alleviate the impact of self-similarity among different fingers. This triangle helps to distinguish each finger from the others so that the network can extract more independent skeleton features. We test our network on three benchmark datasets and demonstrate that our approach achieves performance comparable to state-of-the-art methods with a more lightweight network.

Read the paper · More papers on PaperTik