Speaker normalizing transforms for automatic recognition
Vladimir. Sejnoha, Paul G. Mermelstein · The Journal of the Acoustical Society of America · 1983
A method for the reduction of inter-speaker differences by spectral transformation will be described. Transforms were designed for the additive compensation and frequency scaling of speech spectra based on the log of the energies of 20 channels spaced on the mel-scale. The frequency scaling was implemented as either a linear scaling, or as a nonlinear frequency warp found by dynamic programming. The transforms were then combined so that the additive component principally addressed the differences in the tilt of the spectrum while the frequency scaling mainly treated the dissimilarities in spectrum resonance locations. The transforms were applied to a data base of average spectra of ten steady-state voiced-speech segments for seven male and six female speakers. Both the additive compensation and frequency warping, and particularly the integration of the two, was useful for both within- and across-sex normalization. The overall error rate in recognition experiments on the segment data base was reduced by 50% across sex, 60% within the female speaker group, and by 20% for the male speaker group. The best frequency warp paths were found to be nonlinear and strongly dependent on the speech segment. Parameters for a single, speaker-dependent, combined transform were also derived and the global transform was observed to be as effective as the equivalent speech-segment-dependent transform. The global transform parameters derived from as few as three segments yielded a performance level close to that attained with parameters extracted from the whole set of segments for one speaker. [Research supported by NSERC, Canada.]