Mixed 3D-(2+1)D convolution for action recognition
Bin Yang, ping zhou · 2019
2D CNNS for video-based action modeling ignore the temporal information and treat the multiple frames analogously to channels. In view of this, a mixed convolution structure implemented with ResNet-18 residual network is designed for video feature extracting. The 3D convolution and the (2+1)D convolution are interleaved in sequence throughout the network. Firstly, 2D convolution is performed on input multiple video frames one by one in the spatial. Then, 1D convolution of temporal is performed on the output of 2D convolution. Finally, 3D convolution is performed for spatiotemporal modeling simultaneously. Results show that the mixed convolution structure enhances the transmission of temporal information, improves the ability of video feature extraction and the action recognition accuracy obviously