PTAformer: transformer-based global-local feature interaction framework for behavior recognition
Dongyang Li · Journal of Electronic Imaging · 2025
In the field of computer vision, human action recognition holds significant importance. Although the existing methods based on convolutional neural network (CNN) and Transformer have achieved remarkable results in action recognition tasks, these methods still have deficiencies in the fusion of global and local information. To overcome these limitations, we propose PTAformer—a global–local feature interaction framework based on Transformer for video action recognition. PTAformer combines CNN and BiLSTM layers to extract the underlying information of video frames. In the encoder part, it employs a new deep convolutional module named the improved depthwise convolutional module and an attention mechanism module called the PTA module. Furthermore, we introduce a cross-encoder fusion module to fuse the local features extracted by ResNet and the global features extracted by the PTAformer encoder into a unified feature representation, thereby enhancing the model’s local and global perception capabilities. Innovations have also been made in the normalized part and the classification block to improve the model’s inference speed and performance. In addition, we have created a new dataset named SubVD, which contains abnormal behaviors in the subway. Three key challenges of this dataset are also presented. The proposed model has achieved favorable experimental results on both the self-made dataset SubVD and multiple public datasets. For example, in the self-made dataset SubVD, the Top-1 accuracy is 3.7% higher than that of the ViViT model. This indicates that PTAformer has effectively addressed the challenges proposed in SubVD and has the potential to be applied in practical scenarios.