A Spatiotemporal Decoupled Framework Incorporating Contrastive Learning for RGB-D Gesture Recognition
Yutong Hu, Shaohua Wang, Qingwu Hu, Jiayuan Li · IEEE Sensors Journal · 2025
Gesture recognition is a crucial task in computer vision, with wide-ranging applications in cultural relic exhibitions, smart homes, and virtual reality. Given the persistent and intricate nature of hand movements, achieving accurate gesture recognition typically necessitates the exploration of multimodal cues derived from time-series RGB-D image data. Prior studies have employed coupled dual-branch networks to extract multimodal features; however, they continue to encounter suboptimal challenges in several areas: (1) Tightly coupled modeling methods result in spatiotemporal entanglement, complicating the capture of long-distance dependencies across image sequences, and (2) the inherent heterogeneity between RGB and depth data complicates feature fusion. To address the aforementioned challenges, we introduce a spatio-temporal decoupled framework incorporating contrastive learning for RGB-D gesture recognition, termed DCLNet. This framework aims to improve RGB-D-based gesture recognition from both feature extraction and optimization strategy perspectives. First, we design a decoupled spatiotemporal encoder to acquire dimensionally independent, high-quality features. Second, we incorporate a representation learning stage prior to classification to enhance network optimization. In this stage, a supervised contrastive loss is introduced to constrain the aggregation of similar samples and the separation of heterogeneous samples, facilitating the learning of discriminative and invariant features across modalities. Finally, we implement a self-distillation mechanism to reconstruct correlations among spatiotemporal features. The seamless integration of these innovative designs culminates in a robust multimodal representation, outperforming state-of-the-art methods on two public RGB-D gesture datasets. For example, our DCLNet achieves an accuracy improvement of 1.76% on the IsoGD dataset.