Global spectrum vowel recognition and human performance
Maxine Eskanazi · The Journal of the Acoustical Society of America · 1984
We have shown [Eskenazi and Lienard, J. Acoust. Soc. Am. Suppl. 1 73, S87 (1983)] that global characterizations of the French oral and nasal vowels in a speaker-independent automatic recognition task give generally better recognition results than formant-based methods. In particular, a very rough representation in the frequency domain, characterizing the curvature of the spectrum gave good results using very little reference information for each vowel. By now, changing the analysis that the curvature characterization is based on from an FFT to an LPC has significantly improved global results. This is very close to human intelligibility of the same databases used, in terms of distance between confusion matrices. We shall compare results of human and automatic recognition in order to better evaluate machine performance. There follows a comparison between the FFT and LPC results in order to estimate the pertinence of the information furnished by each in view of vowel recognition and in the light of the problems inherent in a speaker-independent task as well as in specific vowel properties.