Probabilistic Graphical Model for Auto-Annotation, Content-Based Retrieval, and Classification of TV Clips Containing Audio, Video, and Text
Duangmanee Putthividhya, Hagai T. Attias, Srikantan S. Nagarajan, T. W. Lee · 2007
We present a probabilistic graphical model that learns the joint statistical structures of text, audio, and video for the purpose of classification and retrieval of multimedia documents. The proposed model, which we call multi-modal LDA (MM-LDA), builds on the basic latent Dirichlet allocation (LDA) model by postulating common hidden factors, termed topics, that are shared among the 3 data modalities. These hidden topics correspond to patterns of word co-occurrences in multimedia documents and describe how text words co-occur with certain visual and acoustic features. We demonstrate the power of MM-LDA in representing TV clips containing closed-captions, video, and audio, and show promising results in 3 challenging applications: TV clip classification, retrieval, and auto-annotation.