Multi-modal GM-plsa and its application to video classification
Cencen Zhong, Zhenjiang Miao · 2013
To extend standard probabilistic Latent Semantic Analysis (pLSA) to handle continuous quantity, pLSA with Gaussian Mixtures (GM-pLSA) has been proposed, which models the continuous features of terms via a Gaussian Mixture Model (GMM). Stemming from GM-pLSA, this paper presents a multi-modal GM-pLSA (MMGM-pLSA) model to deal with the situation where continuous features from multiple modalities are extracted from one term. Based on our assumption that the multi-modal features of one term independently come from the same latent aspect, multiple GMMs are introduced with each of them depicting the feature distribution of each modality. By doing so, the characteristic of each modality is captured and embodied. To evaluate the performance, a prototype of typical video classification is devised, in which each video clip is interpreted as one document and its sub-shots as terms. Experimental comparisons with other approaches demonstrate the effectiveness of MMGM-pLSA.