Moving Human Focus Inference Model for Action Recognition
Dongli Wang, Hexue Xiao, Fang Ou, Yan Zhou · 2019
Spatio-temporal information is crucial to human action recognition. Inspired by the primary visual cortex of the human brain, a new end-to-end action recognition method based on deep learning was proposed to extract pure spatiotemporal information in videos. It is named moving human focus inference model that can capture long-term temporal information with low time costs and eliminates the interference of dynamic background. It obtains the temporal and spatial information in the video from the temporal pathway with the focus block and spatial pathway, respectively. Then the temporal and spatial information are merged as spatio-temporal information by fusion region. Finally, the softmax layer is used to recognize the actions, followed by the bayes inference block that is used to correct the recognition results through the methods of maximum posterior probability inference. The model is evaluated on two challenge datasets: UCF101 and HMDB51. The experimental results show that the approach has better accuracy and generalization ability compared to the state-of-the-art methods.