Performance analysis of mandarin whispered speech recognition based on normal speech training model
Chen Xueqin, Zhao Heming, Xiaohe Fan · 2016
The magnitude of the decline in performance is very alarming when the features commonly used in normal speech recognition system are directly used as the input feature of whispered speech in the speech recognition system trained by normal speech. In this paper, in order to finding the characteristics of better matching degree between normal and whispered speech, we propose a spectrum sparse-based approach to obtain the feature of speech spectrum structure. We construct a hidden Markov model based speech recognition baseline system to compare the performance of different features. Experimental results show that the proposed feature can perform better on whispered speech recognition based on the speech recognition system trained by normal speech. This means that the recommended feature is better able to express the similarity between the normal speech and the whisper in the spectral topology structure.