Temporal Bottleneck Attention for Video Recognition
Schubert R. Carvalho, Nicolas M. Bertagnolli, Tyler Folkman · 2021 20th IEEE International Conference on Machine Learning and Applications (ICMLA) · 2021
This work introduces temporal bottleneck (TBoT) attention for video-based human action recognition (VHAR). TBoT is a mechanism that aims to build short key-frame sequences from long videos and create more valuable representations for convolution-based models. These compact video clips demonstrate two benefits. First, video recognition models can benefit from compact inputs by learning and modeling the data distribution more quickly and accurately. Second, a network trained on short but informative clips can take advantage of longer sequences when predicting human actions which can improve recognition accuracy. We also propose a multi-head mechanism combining pooling and residual self-attentions that can integrate easily on top of a 2D CNN backbone (ResNet50) and build compelling contexts for human action recognition. We assess the performance of the TBoTNet on several ablation experiments on the Kinetics-400 dataset and compared our results with state-of-the-art architectures on RGB inputs.