Spatio–Temporal Bidirectional Gated Graph Convolutional Network for Skeleton Action Recognition in Dynamic Complex Environments
Lifeng Yin, Yunhan Wang, Qianxi Zhou, Miao Wang, Maohua Sun, Wu Deng · IEEE Internet of Things Journal · 2025
Graph convolutional networks (GCNs) have achieved remarkable success in skeleton-based action recognition. However, most existing studies mainly rely on historical and current data in Internet of Things (IoT) systems, failing to fully exploit the potential of future information. To address this issue, an innovative spatio-temporal bidirectional gated graph convolutional network, namely STBiG-GCN, is designed to enhance the accuracy of action recognition by integrating future information in dynamic and complex environments. Firstly, STBiG-GCN employs bidirectional convolution to enhance the representation of time series features, effectively capturing the dependencies between past and future actions. Secondly, a gating mechanism is designed to dynamically adjust the output, selectively focusing on the most critical features. Finally, by combining skip connections and an efficient multi-scale attention (EMA) mechanism, feature fusion and information flow are optimized, effectively alleviating the vanishing gradient problem in deep networks. On the NTU RGB+D dataset with X-sub and X-view partition standards, STBiG-GCN achieves accuracies of 84.6% and 91.7%, respectively, which are 3.1% and 3.4% higher than those of the ST-GCN model. On the Kinetics-Skeleton dataset, the Top-1 and Top-5 accuracies reach 32.9% and 55.2%, respectively. Additionally, to verify the robustness and generalization performance of the model under data scarcity conditions, experiments were conducted on the Northwestern-UCLA small sample dataset, achieving an accuracy of 92.7%, outperforming existing methods. These significant improvements demonstrate the superiority of STBiG-GCN in capturing action features and provide new solutions for time series analysis.