Video event recognition leveraging hierarchy of semantic concepts

Mohammad Reza Khalifeh Soltanian, Shahrokh Ghaemmaghami · 2017

A new method for exploiting the semantic hierarchical structure of visual concepts in video event recognition task is proposed in this paper. The visual concepts are detected using the readily available Convolutional Neural Network (CNN) structures which make the recognition system extremely efficient in cases with limited hardware resources. The employed CNNs assign scores to each of the predetermined visual concepts in each video frame and the resulting concept scores are fed to the proposed hierarchical post-processing scheme. Our post-processing module takes advantage of the semantic hierarchy of the concepts to enhance the recognition accuracy of event recognition. The hierarchical post-processing works based on the relative shortest distance of concepts specified in Wordnet concept tree and results in a tangible alleviation of uncertainty of the concept scores at the CNN output. The post-processed scores are then delivered to the fine-tuned support vector machine (SVM) classifier to discriminate between the visual event classes. The proposed scheme improves the event recognition accuracy in terms of mean Average Precision (mAP) as demonstrated by the experiments on Columbia Consumer Video (CCV) dataset.

Read the paper · More papers on PaperTik