Vowel identification: Are formants really necessary?

Amir J. Jagharghi, Stephen A. Zahorian · The Journal of the Acoustical Society of America · 1987

It has generally been assumed at least since the time of the comprehensive study by Peterson and Barney [J. Acoust. Soc. Am. 24, 175–184 (1952)] that the formant locations in vowel spectra are the most significant cues to vowel identity. In this experiment vowel spectra were represented by two methods: (A) by the locations of the first three formants, and (B) by the overall smoothed spectral shape in terms of a discrete cosine transform of the power spectra. Stimuli consisted of four repetitions of the widely separated vowels /u/, /i/, /a/, spoken by each of 12 female and 12 male speakers (4⋅24⋅3 = 288 stimuli total). For each of the two spectral encoding methods, A and B, the vowel data were projected to a three-dimensional space such that the vowel categories would be well separated and the vowels within each category well clustered [S. A. Zahorian and A. J. Jagharghi, J. Acoust. Soc. Am. Suppl. 1 79, S8 (1986)]. Significantly better clustering was obtained with method B, based on overall spectral shape, than for method A, based only on the first three formant frequencies. Since these results are not based on perceptual experiments, no direct conclusion can be drawn regarding the perceptual importance of spectral peaks versus overall spectral shape for human perception of vowels. However, the results do indicate that automatic machine identification of vowels can be improved by parameterizing the overall spectral shape rather than only the spectral peaks.

Read the paper · More papers on PaperTik