MediaEval 2012 Tagging Task: Prediction based on One-Best List and Confusion Networks
Yangyang Shi, Martha A. Larson, Pascal Wiggers, Catholijn M. Jonker · 2012
In this paper, we describe our participation in the MediaEval 2012 Tagging task, which requires us to predict the genre labels for videos. We use three different types of models: a conventional support vector machine (SVM), probabilistic generative dynamic Bayesian networks (DBN) and probabilistic discriminative conditional random fields (CRF) to classify the videos based on speech transcripts. As the baseline, SVM uses unigram, bigram and trigram features in a bag-of-words strategy. We also apply the DBN and CRF to take advantage of sequence relationship information in the one-best hypothesis. For confusion networks, the possible words at each time slice are applied as the observation attributes for CRF.