MultiHGR: Multi-Task Hand Gesture Recognition with Cross-Modal Wrist-Worn Devices
Mengxia Lyu, Hao Zhou, Kaiwen Guo, Wangqiu Zhou, Xingfa Shen, Yu Gu · 2024
Hand gesture recognition (HGR) is essential for human-machine interaction. Although the existing solutions achieve good performance in specific tasks, they still face challenges when users navigate through different application contexts, i.e., demanding multi-task ability to support newly arrived HGR tasks. In this paper, we propose the first IMU-vision based system hosted on wrist-worn devices to support multi-task HGR, denoted as MultiHGR. The system introduces a novel two-stage training strategy, i.e., task-agnostic stage to align cross-modal features from unlabeled arbitrary gesture through contrastive learning, and task-related stage to learn modality contributions with limited labeled data in specific tasks through self-attention mechanism. Since only the second task-related stage should be executed for each new task, MultiHGR could accommodate multiple tasks with significant reduced training cost and storage requirement. The evaluation results on three HGR tasks demonstrates that MultiHGR reduces 64.92% training time, and 24.04% storage as compared with traditional multimodal single-task models, and MultiHGR outperforms unimodal single-task models with 14.37%, 19.28%, and 31% improvements in these three tasks, respectively. As compared with state-of-the-art multimodal single-task model, MultiHGR achieves average 6.35% accuracy improvement, along with 65.74% training time reduction.