speaKer Normalization and Adaptation Using

Second-Order Connectionist Networks · 1993

A method for speaker normalization and adaption using connectionist networks is developed. A speaker-specific linear transformation of observations of the speech signal is computed using second-order network units. Classification is accomplished by a multilayer feedforward network that oper- ates on the normalized speech data. The network is adapted for a new talker by modifying the transformation parameters while leaving the classifier fixed. This is accomplished by back- propagating classification error through the classifier to the second-order transformation units. This method was evaluated for the classification of ten vowels for 76 speakers using the first two formant values of the Peterson-Barney data. A classifier optimized on unnormalized data led to a recognition accuracy of 783%. When adapted to each speaker, the accuracy improved to 95.3 %. Another classifier, optimized on linearly normalized data, resulted in a recognition accuracy of 93.2%. When adapted for each speaker from initial transformation parameters estimated by various methods, the accuracy improved to 96.6%. When the speaker-dependent transformation and nonlinear classifier were simultaneously optimized, a vowel recognition accuracy of as high as 97.5% was obtained. The results suggest that rapid speaker adaptation resulting in high classification accuracy can be accomplished by this method.

Read the paper · More papers on PaperTik