Pointing Gesture Interface with Fovea-Lens Camera by Deep Learning using Skeletal Information

Takahiro Ikeda, Satoshi Ueki, Hironao YAMADA · 2023 IEEE/SICE International Symposium on System Integration (SII) · 2023

This paper describes a pointing gesture interface based on deep learning using RGB images acquired by a fovea lens camera as input. This interface consists of hand gesture recognition to determine if the operator is pointing and pointing position estimation. The interface first uses OpenPose to estimate the operator's skeleton from images acquired with the fovea lens camera, which mimics the characteristics of human vision. Next, a hand image is cropped based on the estimated wrist coordinates, and hand gesture recognition is performed. Then, the system classifies which of the pre-defined positions the operator is pointing to by deep learning using the estimated joint positions as input. Evaluation experiments showed that hand gesture recognition achieved an accuracy of more than 95 %. The pointing position estimation achieved an accuracy of more than 85 % when the distance between the camera and the operator was 3 m. In addition, the advantages and challenges of using a fovea lens camera compared to using a standard lens camera and a wide-angle lens camera were discussed.

Read the paper · More papers on PaperTik