DTA: Double LSTM with temporal-wise attention network for action recognition
Yangyang Xu, Lei Wang, Jun Sheng Cheng, Haiying Xia, Jianqin Yin · 2017
In this paper, we propose a new architecture for human action recognition by using a convolution neural networks (CNN) and two Long Short-Term Memory(LSTM) networks with temporal-wise attention model. We call this network the Double LSTM with Temporal-wise Attention network (DTA). The features extracted by our model are both spatially and temporally. The attention model can learn which parts in which frames in a video are relevant to the video label and pay more attention on them. We designed a joint optimization layer (JOL) to jointly process two kinds of feature produced by two LSTMs. The proposed networks achieved improved performance on three widely used datasets — the UCF Sports dataset, the UCF11 dataset and the HMDB51 dataset.