A Human Action Recognition Model Inspired by Multiple Scale Temporal Segments Model Fusion
Hailan Kuang, Wentao Liu, Xinhua Liu, Xiaolin Ma · 2019
Deep convolutional networks have made progress in many tasks in vision, such as image classification, target detection, target segmentation, etc., two-stream ConvNets, C2D+LSTM method and 3DCNN method, successfully introduces the advantages of convolutional neural networks in still image tasks into the field of video action recognition Most of the existing long video recognition methods use fixed sampling as input, which will lose important information of sampling interval. This paper aims to design a spatiotemporal segment with different scales. The structure of the network model is used to learn the video-level descriptions of different kinds of actions. Through this multi-model design, the purpose of optimizing the action classification results with different information distribution is achieved. At the same time, we improved the network structure and optical flow extraction method, designed the network structure using dense block and adopted the TV-Net method to extract the optical flow representation, which also improved the overall model effect. We designed the multi-scale two-stream ConvNets model fusion, the algorithm achieved a satisfactory 95.7% accuracy on the UCF101 data set.