Linear transformations for vowel normalization
Stephen A. Zahorian, Amir J. Jagharghi · The Journal of the Acoustical Society of America · 1989
The results of an evaluation of multivariable linear regression techniques for speaker normalization of vowel data for 11 vowel classes will be presented. The database for the study consisted of the central vowel portions of 2922 CVC syllables obtained from ten males, ten females, and ten children. Each stimulus was represented both by three formants and in terms of overall spectral shape, via the discrete cosine transform coefficients (DCTCs) of the magnitude spectra. In all classification experiments, half the database was used to train the classifier and the other half was used for evaluation. For the case of formants, the classification accuracy on the evaluation data was 63.3% if different speakers were used for training and testing (and thus no speaker normalization), 63.9% if the same speakers were used in the training and test sets but without explicit normalization, and 75.2% with speaker-normalized data. The rates for the corresponding conditions, but with DCTCs as parameters, were 58.0%, 66.1%, and 77.9%. These results indicate a reduction in classification error rate due to speaker normalization of 32.4% for the case of formants and 47.3% for the case of DCTCs. The speaker-normalization techniques are an extension of the techniques reported in the literature [S. F. Disner, J. Acoust. Soc. Am. 67, 253–261 (1980)]. These results also extend our results reported at the Fall 1987 ASA Meeting [S. A. Zahorian and A. J. Jagharghi, J. Acoust. Soc. Am. Suppl. 1 82, S37 (1987)] to a larger, more varied database and indicate that automatic classification of speaker-normalized vowels is comparable to either spectral shape parameters or formants as initial features. [Work supported by NSF.]