Robust Distant Speech Recognition based on Position Dependent CMN
Longbiao Wang, Norihide Kitaoka, Seiichi Nakagawa · 2004
In a distant environment, channel distortion may drastically degrade speech recognition performances. In this paper, we propose a robust multiple microphone speech processing ap-proach based on position dependent Cepstral Mean Normal-ization (CMN). In the training stage, the system measures the transmission characteristics according to the speaker positions from some grid points in the room and estimated the compensa-tion parameters a priori. In the recognition stage, the system es-timates the speaker position and adopts the estimated compen-sation parameters corresponding to the estimated position, and then the system applies the CMN to the speech and performs speech recognition for each microphone. Finally, the maximum vote or the maximum summation likelihood of whole channels (that is, multiple microphones) is used to obtain the final re-sult. In our proposed method, we use utterances emitted from a loudspeaker located at various positions to estimate compensa-tion parameters for a convenient sake, and we also compensate the mismatch between the cepstral means of utterances spoken by human and those emitted from the loudspeaker. Our exper-iments showed that the proposed method improved the perfor-mances of speech recognition system in a distant environment efficiently and it could also compensate the mismatch between voices from human and loudspeaker well. 1.