Attention Based End-to-End Network for Short Video Classification
Hui Zhu, Chao Zou, Zhenyu Wang, Kai Xu, Zihao Huang · 2022
It has been proved that three-dimensional (3D) convolutional kernel can effectively capture local features in the spatiotemporal range of videos, leading to impressive results of various models in video-related tasks. With the introduction of Transformer and the rise of self-attention mechanism, more self-attention models have been used on video representation learning recently. However, there exist limitations of local perception and self-attention operations respectively in both two types of models. Inspired by the global context network (GCNet), we take advantages of both 3D convolution and self-attention mechanism to design a novel operator called the GC-Conv block. The block performs local feature extraction and global context modeling with channel-level concatenation similarly to the dense connectivity pattern in DenseNet, which maintains the lightweight property at the same time. Furthermore, we apply it for multiple layers of our proposed end-to-end network in short video classification task while the temporal dependency is captured via dilated convolutions and bidirectional GRU for better representation. Finally, our model outperforms both state-of-the-art convolutional models and self-attention models on three human action recognition datasets with considerably fewer parameters, which demonstrates the effectiveness.