Audio-visual event classification via spatial-temporal-audio words

Yu Cao, Sung Baang, Shih-Hsi “Alex” Liu, Ming Li, Sanqing Hu · Proceedings - International Conference on Pattern Recognition/Proceedings/International Conference on Pattern Recognition · 2008

In this paper, we propose a generative model-based approach for audio-visual event classification. This approach is based on a new unsupervised learning method using an extended probabilistic latent semantic analysis (pLSA) model. We represent each video clip as a collection of spatial-temporal-audio words, which are generated by fusing the visual and audio features using the pLSA model. Each audio-visual event class is treated as the latent topic in this model. The probability distributions of the spatial-temporal-audio words are learnt from training examples, which include a sequence of videos that represent different types of audio-visual events. Experimental results show the effectiveness of the proposed approach.

Read the paper · More papers on PaperTik