A study of degraded-speech identification based on spectral centroid
Takayuki Furoh, Takahiro Fukumori, Masato Nakayama, Takanobu Nishiura · 2014
Hands-free speech interfaces are developed with the progress of speech recognition techniques. In the conventional automatic speech recognition (ASR) system, normal speech can be recognized with high accuracy. However, the ASR performance is degraded because the human speech is distorted by noise and speaking styles in noisy environments and crisis situations. This problem can be solved by applying suitable acoustic model corresponding to degraded speech. Therefore, we had previously proposed the identification for degradedspeech based on the fundamental frequency (F0), 2nd-order mel-frequency cepstral coefficient (MFCC) and rahmonic. The conventional method can identify normal speech, Lombard speech and shout speech, but it has an insufficient identification performance. This is because the conventional method utilizes acoustic features which are similar in Lombard speech and shout speech. In this paper, we therefore propose degraded-speech identification method based on the spectral centroid, F0, 2nd-order MFCC and rahmonic. The spectral centroid can represent the formant shift to the high-frequency spectrum. In the proposed method, the spectral centroid is utilized for identifying Lombard speech and shout speech. As a result of objective evaluation experiments, we confirmed the effectiveness of the proposed method towards the identification for degraded-speech.