BBN VISER TRECVID 2014 Multimedia Event Detection and Multimedia Event Recounting Systems.
Florian Luisier, Manasvi Tickoo, Walter D. Andrews, Guangnan Ye, Dong Liu, Shih‐Fu Chang, Ruslan Salakhutdinov, Vlad I. Morariu, L. Taylor Davis, Abhinav Kumar Gupta, Ismail Haritaoğlu, Sadiye Guler, Ashutosh Morde · 2014
In this paper, we describe the Raytheon BBN Technologies (BBN) led VISER system for the TRECVID 2014 Multimedia Event Detection (MED) and Recounting (MER) tasks. We present a comprehensive analysis of the different modules: (1) Metadata Generator (MG) – a large suite of audio-visual low-level and sematic features; a set of deep convolutional neural network (DCNN) features trained on the ImageNet dataset; automatic speech recognition (ASR); videotext detection and recognition (OCR). For the low-level features, we used D-SIFT, Opponent SIFT, dense trajectories (HOG+HOF+MBH), MFCC and Fisher Vector (FV) representation. For the semantic concepts, we have trained 1,800 weakly supervised concepts from the Research Set videos and a set of YouTube videos. These concepts include objects, actions, scenes, as well as noun-verb bigrams. We also consider the output layer of the DCNN as a 1,000-dimensional semantic feature. For the speech and videotext content, we leveraged rich confidence-weighted keywords and phrases obtained from the BBN ASR and OCR systems. (2) Event Query Generation (EQG)- linear SVM event models are trained for each feature and combined using probabilistic late fusion framework. Our system involves both SVM-based and query-based detections, to achieve superior performance despite the varying number of positive videos in the different training conditions. We present a thorough study and evaluation of different