MSAHTA: Mixed Spatial Attention and Hierarchical Temporal Aggregation for Action Recognition
Jinyuan Feng, Dan Yang, Yongxin Ge, Xiaolei Qin, Yida Chen, Yuangan Wang · 2019
Spatial-temporal feature modeling in videos is a fundamental yet knotty problem in computer vision. In this paper, we present a novel model, termed as Mixed Spatial Attention and Hierarchical Temporal Aggregation (MSAHTA), to highlight key spatial feature and aggregate temporal information in the multi-scale and hierarchical manner. MSAHTA is constructed by Spatial Attention module (SA) and Hierarchical Temporal Aggregation module (HTA). Specifically, the SA module aims to enhance the discrimination of spatial-feature by soft attention, while the HTA module aggregates the enhanced spatial-feature of each frame over time by means of a set of hierarchical 3D convolution operation. Furthermore, MSAHTA is an extensible module that can be used as a plugin for the existing 2D CNN. Our method obtains state-of-the-art results on two widely applicable datasets: UCF101 and HMDB51.