A semi-Markov model for speech segmentation with an utterance-break prior
Mark Sinclair, Peter Bell, Alexandra Birch, Fergus R. McInnes · 2014
Speech segmentation is the problem of finding the end points of a speech utterance for passing to an automatic speech recogni-tion (ASR) system. The quality of this segmentation can have a large impact on the accuracy of the ASR system; in this pa-per we demonstrate that it can have an even larger impact on downstream natural language processing tasks – in this case, machine translation. We develop a novel semi-Markov model which allows the segmentation of audio streams into speech ut-terances which are optimised for the desired distribution of sen-tence lengths for the target domain. We compare this with exist-ing state-of-the-art methods and show that it is able to achieve not only improved ASR performance, but also to yield signifi-cant benefits to a speech translation task. Index Terms: speech activity detection, speech segmentation, machine translation, speech recognition