Action Recognition by Composite Deep Learning Architecture I3D-DenseLSTM
Yuri Yudhaswana Joefrie, Masaki Aono · 2019
Two-Stream Convolutional Neural Networks have shown remarkable results for action recognition in videos. In this paper, we adopt the two-stream principle to exploit appearance and motion features from a video. We extend two-stream to three-stream to achieve diversity. Two streams are employed to capture spatial and temporal aspects, while another stream is used to capture motion modality. In a pre-processing step, we prepare the data, from where we extract features and optical flow frames. We conducted experiments on Moment-In-Time dataset available publicly. Our novel network architecture is composed of three main parts: DenseLSTM component (DenseNet-like skip connections plus LSTM), a single LSTM component, and the Inflated 3D. Altogether we call our model I3D-DenseLSTM. Through experiments, we demonstrate that our proposed model outperforms several baseline models.