Enhancing dynamic gesture recognition and human–computer interaction through integrated GCN–transformer architecture with transfer learning

Minna Liu, Xinming Guo · Alexandria Engineering Journal · 2026

Dynamic gesture recognition is a core technology for human–computer interaction, yet existing methods face prominent limitations in hand spatial topology modeling, long–short-term temporal feature fusion, and generalization in complex practical scenarios. To tackle these challenges, this paper proposes the TL-BGGT model, which integrates sparse graph convolutional GCN for spatial feature extraction and Transformer for temporal feature fusion, and designs a transfer learning strategy with pre-training and differential fine-tuning to enhance cross-scenario adaptability. Extensive experiments on LAP, MSR Gesture and HaGRID datasets confirm TL-BGGT’s state-of-the-art performance: it achieves 99.06% accuracy on LAP, 82.15% on MSR Gesture (3.3% higher than MEMP Network), and 99.02% on HaGRID, with mAP at 91.35% and 96.94% on LAP and HaGRID, outperforming mainstream baselines. The model also exhibits excellent convergence stability and real-scene generalization, offering a reliable technical support for practical human–computer interaction systems. Finally, its limitations are analyzed, and future directions including multi-modal fusion, domain adaptation optimization, and lightweight design are outlined.

Read the paper · More papers on PaperTik