Robust distant speaker recognition based on position dependent cepstral mean normalization
Longbiao Wang, Norihide Kitaoka, Seiichi Nakagawa · 2005
In a distant environment, channel distortion may drastically degrade speaker recognition performance. In this paper, we propose a robust speaker recognition method based on posi-tion dependent Cepstral Mean Normalization (CMN) to com-pensate the channel distortion depending on the speaker posi-tion. It is shown in [1] that the position dependent CMN is robust for speech recognition in a distant environment. We ex-tend this method to the speaker recognition and show that this method is much effective to speaker recognition. In the train-ing stage, the system measures the transmission characteristics according to the speaker positions from some grid points to the microphone in the room and estimated the compensation pa-rameters a priori. In the recognition stage, the system esti-mates the speaker position and adopts the estimated compen-sation parameters corresponding to the estimated position, and then the system applies the CMN to the speech and performs speaker recognition. In our past study, we proposed a new text-independent speaker recognition method by combining speaker-specific Gaussian Mixture Models (GMMs) with syllable-based HMMs adapted to the speakers by MAP [2]. The robustness of this speaker recognition method for the change of the speaking style in close-talking environment was evaluated in [2]. We in-tegrated this method to the proposed position dependent CMN for distant speaker recognition. Our experiments showed that the proposed method improved the speaker recognition perfor-mance remarkably in a distant environment. 1.