Speaker-independent automatic vowel recognition based on overall spectral shape versus formants
Stephen A. Zahorian, Amir J. Jagharghi · The Journal of the Acoustical Society of America · 1987
Automatic recognition experiments were performed to compare overall spectral shape versus formants as speaker-independent acoustic parameters for vowel identity. Stimuli consisted of four repetitions of 11 vowels spoken by 17 female speakers and 12 male speakers (29*11*4 = 1276 total stimuli). Formants were computed automatically by peak picking of 12th-order LP model spectra. Spectral shape was represented using three methods: (1) by a cosine basis vector expansion of the power spectrum: (2) as the output of a 16-channel, 1/3-oct filter bank; and (3) as the output of a 16-channel mel-spaced filter bank. Automatic recognition was based on maximum likelihood estimation in a multidimensional space. For all cases considered, the representations based on spectal shape resulted in significantly higher recognition accuracy than for recognition based on only three formants. For example, using the entire database of all speakers and 11 vowels, recognition based on spectral shape was about 85% vs 69% for three formants. If the data were restricted to female speakers and the seven vowels /a,i,u,æ,ɝ,ɪ,ɛ/, recognition was about 97% based on spectral shape versus 84% for formants. These results indicate that, at least for automatic recognition of vowels, spectral peak detection is neither necessary nor sufficient. [Work supported by NSF.]