The Multimedia Understanding Group at TRECVID-2010

Christos Diou, George Stephanopoulos, Anastasios N. Delopoulos · 2010

This is a report of the Multimedia Understanding Group participation in TRECVID-2010, where we submitted full runs for the Semantic Indexing (SIN) task. Our submission aims at experimentally evaluating three research items, that are important for work that is currently in progress. First, we examine the use of bag-of-words audio features for video concept detection, with noisy and/or low-quality video data. Although audio is important for some concepts and has shown promising results at other datasets, the results indicate that it can also lead to a decrease in performance when the quality is low and the negative examples are not adequately represented. We also explore the possibility of using a cross-domain concept fusion approach for reducing the number of dimensions at the final classifier. The corresponding experiments show, however, that when drastically reducing the number of dimensions the effectiveness drops. Finally, we also examined a transformation of the feature space, using a set of functions that are parametrically constructed from the data. Semantic Indexing Runs 1. F A MUG-AUTH 1: Early fusion of all available low-level features. 2. F A MUG-AUTH 2: Early fusion of all available low-level features except audio. 3. F A MUG-AUTH 3: Feature space transformation based on a sample of the training data. 4. F C MUG-AUTH 4: Concept fusion using a limited number of base classifiers. 1

Read the paper · More papers on PaperTik