TokyoTech+Canon at TRECVID 2011
Nakamasa Inoue, Kotaro Mori, Zhuolin Liang, Mengxi Lin, Koichi Shinoda, Shunsuke Sato · 2011
The aim of this section is to develop a high-performance semantic indexing system using Gaussian mixture model (GMM) supervectors and tree-structured GMMs [1, 2]. GMM spervectors corresponding to six types of audio and visual features are extracted from video shots by using tree-structured GMMs. The computational cost of maximum a posteriori (MAP) adaptation for estimating GMM parameters are reduced by tree-structured GMMs by keeping accuracy at high levels. Our best result was 17.3 % in terms of Mean InfAP, which was ranked 1st over all semantic indexing runs in the full task. 1.1 Feature extraction The following six types of visual and audio features are extracted from video data: 1. SIFT features with Harris-Affine detector (SIFT-Har) The scale invariant feature transform (SIFT) proposed by Lowe [3] is a local feature extraction method that is widely used for object categorization since it is invariant to image scaling and changing illumination. The Harris-Affine detector [5], which is an extension of the Harris corner detector, improves robustness against affine transform of local regions. These features are extracted from every other frame, and principal component analysis (PCA) is applied to reduce their dimen-sions from 128 to 32.