University of Amsterdam at THUMOS 2015
Mihir Jain, Jan C. van Gemert, Pascal Mettes, Cees G. M. Snoek · 2015
This notebook paper describes our approach for the action classification task of the THUMOS 2015 benchmark challenge. We use two types of representations to capture motion and appearance. For a local motion description we employ HOG, HOF and MBH features, computed along the improved dense trajectories. The motion features are encoded into a fixed-length representation using Fisher vectors. For the appearance features, we employ a pre-trained GoogLeNet convolutional network on video frames. VLAD is used to encode the appearance features into a fixed-length representation. All actions are classified with a one-vs-rest linear SVM.