Improved DNN-based segmentation for multi-genre broadcast audio

Lin Wang, Chengyue Zhang, Philip C. Woodland, Mark Gales, Penny Karanasou, Pierre Lanchantin, Xiaobing Liu, Yanmin Qian · 2016

Automatic segmentation is a crucial initial processing step for processing multi-genre broadcast (MGB) audio. It is very challenging since the data exhibits a wide range of both speech types and background conditions with many types of non-speech audio. This paper describes a segmentation system for multi-genre broadcast audio with deep neural network (DNN) based speech/non-speech detection. A further stage of change-point detection and clustering is used to obtain homogeneous segments. Suitable DNN inputs, context window sizes and architectures are studied with a series of experiments using a large corpus of MGB television audio. For MGB transcription, the improved segmenter yields roughly half the increase in word error rate, over manual segmentation, compared to the baseline DNN segmenter supplied for the 2015 ASRU MGB challenge.

Read the paper · More papers on PaperTik