Fusion Based Deep CNN for Improved Large-Scale Image Action Recognition
Yukhe Lavinia, Holly H. Vo, Abhishek Verma · 2016
Still image-based action recognition is a process of labeling actions captured in still images. We propose a fusion method that concatenates two and three deep convolutional neural networks (CNN). After examining the classification accuracy of each deep CNN candidates, we inserted a 100 dimensional fully-connected layer and extracted features from the new 100 dimensional and the last fully-connected layers to create a pool of candidate layers. We form our fusion models by concatenating two or three layers from this pool—one from each model—and trained and tested them on the large-scale Stanford 40 Actions still image dataset. We forwarded the concatenated features to Random Forest and SVM for classification. Our experiments show that our fusion of two deep CNN models achieved better accuracy than the individual models, with the best performing fusion duo of 80.351% accuracy. The fusion of three models increases the accuracy even further, performing better than both the individual and the fusion of two models, with 81.146% accuracy. Moreover, we also investigate the classification difficulty level of the Stanford 40 Actions category.