Optimal automatic speech recognition system selection for noisy environments

Yuuki Tachioka, Tomohiro Narita · 2016

To improve the performance of noisy automatic speech recognition (ASR), it is effective to prepare multiple ASR systems that can address the large varieties of noise. However, the optimal ASR system is different for each environment and mismatches between training and testing degrade ASR performance. In this situation, the overall system combination of multiple systems is effective; however, the computational resources increase in proportion to the number of systems. This paper proposes a method to select an optimal single system from multiple systems. The selection is based on the estimated word error rates of a respective system by using the i-vector similarities between training and test data. The experiments on the third CHiME challenge show that our proposed method can efficiently select a single system from multiple systems with different speech enhancement and feature transformation methods to improve the overall performance without increasing computational resources.

Read the paper · More papers on PaperTik