High-level feature weighted GMM network for audio stream classification

Rongqing Huang, John H. L. Hansen · 2004

The problem of unsupervised audio classification con-tinuous to be a challenging research problem which sig-nificantly impacts ASR and Spoken Document Retrieval (SDR) performance. This paper addresses novel ad-vances in audio classification for speech recognition. A new algorithm is proposed for audio classification, which is based on Weighted GMM Network (WGN). Two new high-level features: VSF (Variance of the Spectrum Flux) and VZCR (Variance of the Zero-Crossing Rate) are used to pre-classify the audio and supply weights to the output probabilities of the GMM networks. The classification is then implemented using weighted GMM networks. Eval-uations on a standard data set — DARPA Hub4 Broadcast News 1997 evaluation data, shows that the WGN classi-fication algorithm achieves over a 50 % improvement ver-sus the GMM network baseline algorithm. The WGN also obtains very satisfactory results on the more diverse and challenging NGSW (National Gallery of the Spoken Word [8]) corpus. Classification based on segmentation method is also explored. 1.

Read the paper · More papers on PaperTik