Classification of vowels using spectrum moments
Paul Milenkovic, Karen M. Forrest · The Journal of the Acoustical Society of America · 1988
Formant frequencies form a set of acoustic features that can be used to classify vowels phonetically [A. K. Syrdal and H. S. Gopal, J. Acoust. Soc. Am. 79, 1086–1100 (1986)]. The advantage of formant frequencies is that they are unaffected by changes in the speech spectrum slope attributable to the voice source. The disadvantage of formant frequencies is in the large errors in frequency that can result from a misclassification of the formant peaks of the speech spectrum. A set of features based on statistical moments computed from the speech spectrum has been successful in classifying a broad class of obstruent sounds [K. Forrest et al., J. Acoust. Soc. Am. Suppl. 1 82, S84 (1987)]. In the case of vowel sounds, moment analysis eliminates the problem of identifying formant peaks but introduces the problem of variability in the moment features resulting from changes in the voice source. In order to correct for between subject variation in the voice source, the long-term average spectrum of all of the speech samples recorded from each subject is computed and used to normalize vowel spectra prior to computing moments. Between subject classification results for adult subjects are presented. [Work supported by NIH grants NS 21516 and NS 13274.]