Sub-sentence discourse models for conversational speech recognition
K.W. Ma, George Zavaliagkos, Marie W. Meteer · 2002
According to discourse theories in linguistics, conversational utterances possess an informational structure that partitions each sentence into two portions: a given and new. We explore this idea by building sub-sentence discourse language models for conversational speech recognition. The internal sentence structure is captured in statistical language modeling by training multiple n-gram models using the expectation-maximization algorithm on the Switchboard corpus. The resulting model contributes to a 30% reduction in language model perplexity and a small gain in word error rate.