L1-regularized Logistic Regression Stacking and Transductive CRF Smoothing for Action Recognition in Video
Svebor Karaman, Lorenzo Seidenari, Andrew D. Bagdanov, Alberto Del Bimbo · 2013
For the 2013 THUMOS challenge we built a bag-of-features pipeline based on a variety of features extracted from both video and keyframe modalities. In addition to the quantized, hard-assigned features provided by the orga-nizers, we extracted local HOG and Motion Boundary His-togram (MBH) descriptors aligned with dense trajectories in video to capture motion. We encode them as Fisher vec-tors. To represent action-specific scene context we compute local SIFT pyramids on grayscale (P-SIFT) and opponent color keyframes (P-OSIFT) extracted as the central frame of each clip. From all these features we built a bag-of-features pipeline using late classifier fusion to combine scores of in-dividual classifier outputs. We further used two complemen-tary techniques that improve on the basic baseline with late fusion. First, we improve accuracy by using L1-regularized logistic regression (L1LRS) for stacking classifier outputs. Second, we show how with a Conditional Random Field (CRF) we can perform transductive labeling of test samples to further improve classification performance. Using our features we improve on those provided by the contest orga-nizers by 8%, and after incorporating L1LRS and the CRF by more than 11%, reaching a final classification accuracy of 85.7%. 1.