A comparison of two models of human vowel recognition
Dawn L. Dutton, James R. Sawusch, Kam-Cheong Tsoi · The Journal of the Acoustical Society of America · 1987
Two models of human vowel recognition were simulated and their performances compared. The first was proposed by Syrdal and Gopal [J. Acoust. Soc. Am. 79, 1086–1100 (1986)] and modified for implementation. Five dimensions composed of critical band differences between spectral peaks (including F0) were computed for each input stimulus. Identification was based on the best match of the input to a set of prototypes. The second was a full octave filtering model simulating central auditory processing. Parallel filters combined all spectral peaks within three critical bands into a single peak. The match between each stimulus and the prototypes was then used as the basis for categorization. The prototypes used in each model were based on Peterson and Barney [J. Acoust. Soc. Am. 24, 175–184 (1952)]. The performance of the two models in identifying vowels will be described. In addition, performance of the models was compared to data of human subjects for tone analogs of vowel and single formant vowels. Relative strengths and weaknesses of the two models will be discussed. [Work supported by NINCDS.]