Automatic Synchronization of live speech and its Transcripts based on a frame-synchronous likelihood ratio test
Jie Gao, Qingwei Zhao, Yonghong Yan · 2010
In this paper, we present our initial efforts in the task of Automatically Synchronizing live spoken Utterances with their Transcripts (textual contents) (ASUT) when the texts are known. We treat it as a online speech-text alignment problem. And it is further simplified into the problem of on-the-fly detecting of the end time of a spoken utterance given its textual content. A general framework called frame-synchronous likelihood ratio test (FS-LRT) procedure is proposed for this end time detection task and explored with the hidden Markov models (HMMs). The property of FS-LRT is studied empirically. Extensive experiments indicate that our proposed approach shows satisfying performance. In addition, FS-LRT has been successfully applied in a subtitling system for live broadcast news.