GENIE TRECVID 2011 Multimedia Event Detection: Late-Fusion Approaches to Combine Multiple Audio-Visual features.
A. G. Amitha Perera, Sangmin Oh, Matthew J. Leotta, Ilseo Kim, Byungki Byun, Chin‐Hui Lee, Scott McCloskey, Jingchen Liu, Ben Miller, Zhi feng Huang, Arash Vahdat, Weilong Yang, Greg Mori, Kevin Tang, Daphne Koller, Li Fei-Fei, Shuo Li, Gang Chen, Jason J. Corso, Yun Fu · 2011
For TRECVID 2011 MED task, the GENIE system incorporated two late-fusion approaches where multiple discriminative base-classi ers are built per feature, then, combined later through discriminative fusion techniques. All of our fusion and base classi ers are formulated as one-vs-all detectors per event class along with threshold estimation capabilities during cross-validation. Total of ve di erent types of features were extracted from data, which include both audio or visual features: HOG3D, Object Bank, Gist, MFCC, and acoustic segment models (ASMs). Features such as HOG3D and MFCC are low-level features while Object Bank and ASMs are more semantic. In our work, event-speci c feature adaptations or manual annotations were deliberately avoided, to establish a strong baseline results. Overall, the results were competitive in the MED11 evaluation, and shows that standard machine learning techniques can yield fairly good results even on a challenging dataset. Summary of Submitted Runs 1. GENIE_MED11_MED11TEST_MEDFull_AutoEAG_p-MFoMaudiovisual_1: Fusion by MFoM with 5 audio-visual features 2. GENIE_MED11_MED11TEST_MEDFull_AutoEAG_c-MLE_1: Fusion by MLEs with 5 audio-visual features 3. GENIE_MED11_MED11TEST_MEDFull_AutoEAG_c-MFoMvideo_1: Identical to #1, with only 3 video features 4. GENIE_MED11_MED11TEST_MEDFull_AutoEAG_c-MFoMaudio_1: Identical to #1, with only 2 audio features 1