Vision Based Human Action Classification Using CNN model with Mode Calculation
Iffat Zabin Tonu, Abdul Matin, Hafsa Binte Kibria · 2021
Video classification is a wide field of study in computer vision research. Many types of approaches have been used for video classification: sensor-based approaches, deep learning-based approaches, handcrafted approaches like hog and blob detection, optical flow-based, sliding window-based, a nd many more. Among all of the strategies, Deep learning has shown quite progressive outcomes. In this work, a CNN-based model has been used. The CNN-based approach for video classification is considered a ladder to the world of deep learning-based computer vision research. The proposed methodology is quite simple. Considering the whole video as a series of frames, it predicts each frame using CNN and then identifies t he c lass of the video using mode calculation of the predictions. To predict the frame, we have followed two types of models. The first approach was building a CNN model from scratch and training the model with our image data, and then using the model for prediction. The second approach used transfer learning for feature extraction and then used those features to train a modified C NN model. The proposed approaches have been tested on a compromised UCF101 sports dataset, and the obtained accuracy was 88.35% for the first model and 93.15% for the second one.