Spatio-Temporal Feature Extraction for Human Action Recognition in Visual Communication Systems Using Improved ST-GCN
Weitao Gao, Chenbin Wang, Changxin Chen, Tiehua Ma, Wenchao Guo · Journal of Circuits Systems and Computers · 2025
Using the Spatio-Temporal Graph Convolutional Networks (ST-GCN) model, this study applies an improved spatial-temporal feature extraction mechanism to enhance the performance of visual communication system in human action recognition for dynamic graphics. First, the Canny operator is used to extract the foreground contour, and the Conditional Random Field (CRF) is used to suppress the background interference information. Then, MediaPipe is used to extract human key points from the video frame to identify the position and connection relationship. Based on ST-GCN, spatial features of action key points are extracted, and the adjacency matrix is defined to describe the connection relationship between key points. At the same time, dilated convolution is applied in the time dimension to capture long-term dependencies by increasing the receptive field of the convolution kernel. Multiple graph convolutional layers and temporal convolutional layers are stacked, and both spatial and temporal features are captured. Finally, the Improved ST-GCN model (I-ST-GCN) is integrated into the visual communication system, and its performance in action recognition in dynamic images is evaluated through multi-scene tests. The results show that the Mean Average Precision (mAP) value of I-ST-GCN in four groups of tests is above 0.94, and the Intersection over Union (IoU) value for Step Down complex occluded action recognition is 0.90. The improved method adopted can realize the precise action recognition of complex dynamic graphics by the visual communication system.