Learn to Recognize Actions Through Neural Networks
Zhenzhong Lan · 2015
This research seeks to develop neural network techniques to effectively recognize actions in videos. The proposed study will lead to a deeper understanding of how neural network algorithms can help AI systems to understand motions. It will also realize an action recognition system that significantly outperforms current state-of-the-art. In addition, we will investigate the extent to which our work can be beneficial to understanding how brains perceive and analyze actions. Recent data-driven neural network approaches such as convolutional neural networks have been successful for object recognition. However, learning to recognize actions has proven to be quite a challenge due to the difficulty of getting enough labels, processing large-scale video data, and capturing motion information from videos. Therefore, we leverage effective techniques from local hand-crafted methods to help neural network algorithms learn motion features. These techniques include learning from video and optical flow volumes that follow motion trajectories, pooling features from videos played at multiple frame rates to achieve speed invariance, extending the local descriptors with normalized locations to incorporate spatial-temporal information, and a training-free re-ranking technique to exploit the relationship among classes. We also discuss a fundamental problem of whether should we learn time-aware models and what our models actually capture when we feed them temporal data. Finally, we discuss the connection of our research with how brains perceive and recognize motions.