IBM Research TRECVID-2008 Video Retrieval System
Apostol Natsev, Matthew L. Hill, John R. Smith, Lexing Xie, Rong Yan, Shenghua Bao, Dong Wang, Michele Merler, Yi Zhang · 2008
In this paper, we describe the IBM Research system for indexing, analysis, and retrieval of video as applied to the TREC-2008 video retrieval benchmark. This year, focus of the system improvement was on large-scale learning, cross-domain detection, and interactive search. A. High-level concept detection: 1. A ibm.Baseline 5: Baseline runs with randomsubspace bagging; 2. A ibm.BaseSSL 4: Fusion of baseline runs and principal component semi-supervised support vector machines(P CS 3 V M); 3. A ibm.BaseSSLText 3: Fusion of A ibm.BaseSSL 4 and text search results; 4. C ibm.CrossDomain 6: Learning on data from web domain; 5. C ibm.BNet 2: Multi-concept learning with baseline, P CS 3 V M and web concepts; 6. C ibm.BOR 1: Best overall runs by compiling the best models based on heldout performance for each concept. Overall, almost all the individual components can improve the mean average precision after fused with the baseline results. To summarize, we have the following observations from our evaluation results: 1) The baseline run using random-subspace bagging offers a reasonable starting performance with a more efficient learning process than standard SVMs; 2) By learning on both feature space and unlabeled data,