Improving language models for ASR using translated in-domain data

Stefan Kombrink, Tomáš Mikolov, Martin Karafiát, Lukáš Burget · 2012

Acquisition of in-domain training data to build speech recognition systems for under-resourced languages can be a costly, time-demanding and tedious process. In this work, we propose the use of machine translation to translate English transcripts of telephone speech into Czech language in order to improve a Czech CTS speech recognition system. The translated transcripts are used as additional language model training data in a scenario where the baseline language model is trained on off- and close-domain data only. We report perplexities, OOV and word error rates and examine different data sets and translators on their suitability for the described task.

Read the paper · More papers on PaperTik