Classification of Music and Speech in Mandarin News Broadcasts

Chuan Liu, Lei Xie, Helen M. L. Meng · 2007

Audio scene analysis refers to the problem of class ifying segments in a continuous audio stream according to content, e.g. speech versus non-speech, music, ambi ent noise, etc. Techniques that support such autom atic segmentation is indispensable for multimedia information processing. For example, it is a precursor to processes such as indexing of speech segments by automatic speech recognition, automatic story segmentation based on recognition transcript s, speaker diarization, etc. This paper describes our work in the developm ent of a speech/music discriminator for Mandarin broadcast news audio. We formed a high-dimensional feature vector that in cludes LPCC, LPS and STFT coefficients totaling 94 in all. We also experimented with three classifiers - the KNN, SVM and MLP. Experiments based on the Voice of America Mandarin news broadcasts show high classification performance with F-measure=0.98. The SVM also strikes the best balance in terms of classification performance and computation time (re al-time) among the three classifiers. 1

Read the paper · More papers on PaperTik