Convolutional Neural Network using Stacked Frames for Video Classification
Itthisak Phueaksri, Sukree Sinthupinyp · 2019
We propose a Convolutional Neural Network model with stacked frame images for video classification. In this research, one of the challenges is that each video is more than 3,600 seconds long. We extracted ordered frames by skipping some frames from each video. Under our assumption, it is not practical to train each video with all the extracted frames as input because it can cause building a huge model. Hence, we created a median filter layer for reducing the number of frames before training. In the experiment, we extracted 500 frame images from each video. In the reduction process with a median filter layer, we were able to reduce the number of frames from 500 frame images to 50 median images. In the training process, we trained our model that used binary classification to classify seven classes of TV programs. With 1,092 video clips of TV programs, our model achieved 79.38% average accuracy.