Robust lecture speech translation for speech misrecognition and its rescoring effect from multiple candidates

Koya Sahashi, Norioki Goto, Hiroshi Seki, Kazumasa Yamamoto, Tomoyoshi Akiba, Seiichi Nakagawa · 2017

We describe a scheme to translate spoken English lectures into Japanese consisting of a deep neural network based English automatic speech recognition system (ASR) and an English to Japanese phrase-based statistical machine translation system (SMT). The bad influence of speech misrecognition for the translation model is focused. For coping with bad influence caused by speech misrecognition, we utilized the actual misrecognition results as a parallel corpus. We prepared four ASR systems. Pairs of the results including speech misrecognition and the correct translation into target language are added to an original parallel corpus. When the sentences including misrecognition were added to an original corpus, the baseline model was improved. Next, we prepared speech misrecognition results by using a simulated ASR system. This method also improved the baseline system by about 2.0 BLEU as well as actual ASR systems. Finally, we investigated the effectiveness of optimal selection from multiple candidates / outputs by rescoring based on language models. We found that the performance was improved further by 2.0 ~ 3.0 BLEU, if we can select the optimal candidate.

Read the paper · More papers on PaperTik