Phonetically adaptive cepstrum mean normalization for acoustic mismatch compensation
Masatoshi Morishima, Toshihiro Isobe, Junichi Takahashi · 2002
We propose a new technique that compensates for an acoustic mismatch. This technique is simple and can estimate the acoustic mismatch more accurately than conventional cepstrum mean normalization (CMN), because it takes into consideration the kind of phonemes and their frequency, and can calculate the acoustic mismatch in detail. In this procedure the acoustic mismatch can be estimated as the difference between the centroid vector of distorted speech and that of acoustic models. The cepstral mean of distorted speech is the centroid vector including the distortion. The centroid vector calculated from parameters of acoustic models is regarded as the centroid vector when the distorted speech is assumed to be clean speech. The acoustic models used for calculation are for phonemes that appear in the transcription of the speech. This technique achieves a high word error reduction rate of 73% for ordinary analog telephone speech and 70% for wireless telephone handset speech.