Long-term 3D Convolutional Fusion Network for Action Recognition
Yanyi Zhang, Kuangrong Hao, Xue‐song Tang, Bing Wei, Lihong Ren · 2019
Applications of deep convolutional networks have achieved success in the field of video action recognition. However, there are still existing problems: 1) Most motion recognition only uses a single video segment as research object, which is difficult to represent the coherent process of the action in whole video because of the lack of ability to use the complementary information; 2) Most studies have not specifically enhanced the ability to characterize movements as time goes on, just put temporal information in the same class as spatial information. A Long-term 3D Convolutional Fusion Network (LT3D-CFN) is proposed to aim at these problems. LT3D-CFN could be summarized as the following two aspects: CNN in two-stream network has been replaced by 3DCNN because of the ability of extracting features from both the spatial and the temporal dimensions of one video clip. In addition, long-term associations between clips of one motion video are established by adding deep LSTM network. The proposed architecture LT3D-CFN, is trained and evaluated on UCF-101 dataset where the proposed architecture achieves superior performers.