Spatial-Temporal Network Model for Action Recognition Based on Keyframe Extraction
Dongjie Zhao, Haili Zhao, Kai Yao, Jiahe Meng, Meihui Liang · 2024
This paper proposes a spatiotemporal network model TransTC3D for behavior recognition. The proposed SIC is used to sample video key frames, and then the key frames are sent to the TransTC3D network to complete behavior recognition. Based on the fine-grained feature extraction capability of the C3D model, the TransTC3D network adds a temporal attention mechanism (SVAM) and a channel attention mechanism (SCAM) to the seventh and twelfth layers respectively. The high-dimensional fine-grained feature information is extracted through the improved C3D network to obtain fine-grained spatial and temporal feature information, which is then sent to the Transformer for long-term sequence modeling to finally achieve behavior recognition. It has been verified that SIC sampling combined with the TransTC3D model has achieved an effect of 90.23% on the UCF101 dataset.