Video Event Recognition using Two-Stream Convolutional Neural Networks

Mohammad Alijanpour, Abolghasem Asadollah Raie · 2021

Event Recognition aims to classify activities done by a group of people and/or objects in a video into different classes. Convolutional Neural Networks (CNNs) have shown promising results in computer vision tasks, specifically human activity recognition. This paper presents a Two-Stream Convolutional Neural Network for video Event Recognition which uses Xception architecture as its backbone. The dual streams of this method are responsible for extracting spatial and temporal features. The first stream extracts spatial features from RGB frames and the second stream extracts motion features from optical flow images. Due to optical flow estimation complexity, another method for motion estimation between two frames called Thresholding Frame Subtraction (TFS) is used which can speed up the preprocessing phase. The method showed recognition accuracies of 85.6% and 73.2% on VIRAT ground dataset in the best case using optical flow and TFS images respectively which outperform other works on this dataset.

Read the paper · More papers on PaperTik