New colour fusion deep learning model for large-scale action recognition

Yukhe Lavinia, Holly H. Vo, Abhishek Verma · International Journal of Computational Vision and Robotics · 2020

In this work we propose a fusion methodology that takes advantage of multiple deep convolutional neural network (CNN) models and two colour spaces RGB and oRGB to improve action recognition performance on still images. We trained our deep CNNs on both the RGB and oRGB colour spaces, extracted and fused all the features, and forwarded them to an SVM for classification. We evaluated our proposed fusion models on the Stanford 40 Action dataset and the People Playing Musical Instruments (PPMI) dataset using two metrics: overall accuracy and mean average precision (mAP). Our results prove to outperform the current state-of-the-arts with 84.24% accuracy and 83.25% mAP on Stanford 40 and 65.94% accuracy and 65.85% mAP on PPMI. Furthermore, we also evaluated the individual class performance on both datasets. The mAP for top 20 individual classes on Stanford 40 lies between 97% and 87%, on PPMI the individual mAP class performance lies between 87% and 34%.

Read the paper · More papers on PaperTik