Latent semantic indexing for semantic content detection of video shots
Fabrice Souvannavong, Bernard Mérialdo, Benoît Huet · 2005
Low-level features are now becoming insufficient to build efficient content-based retrieval systems. The interest of users is not any more to retrieve visually similar content, but they expect retrieval systems to find documents with similar semantic content. Bridging the gap between low-level features and semantic content is a challenging task necessary for future retrieval systems. Latent semantic indexing (LSI) was successfully introduced to efficiently index text documents. We propose to adapt this technique to efficiently represent the visual content of video shots for semantic content detection. Although we restrict our approach to visual features, it can be extended with minor changes to audio and motion features to build a multi-modal system. The semantic content is then detected thanks to two classifiers: k-nearest neighbors and neural network classifiers. Finally, in the experimental section we show the performances of each classifier and the performance gain obtained with LSI features compared to traditional features.