Recoguition of vowels using artificial neural networks
Hong C. Leung, Victor W. Zue · The Journal of the Acoustical Society of America · 1988
This paper is concerned with the application of artificial neural networks to phonetic recognition. This work is motivated by the observation that improved knowledge of feature extraction is often overshadowed by relative ignorance on how to combine them into a robust decision. The goal is to investigate how the self-organizing framework of artificial neural networks can be exploited to enable different acoustic cues to interact. The investigation is couched in experiments that recognize the 16 vowels in American English, using some 10 000 tokens in all phonetic contexts. The tokens were extracted from 1000 sentences spoken by 140 males and 60 females. It was found that, by replacing the mean-squared error metric with a weighted one to train a multilayer perception, better recognition accuracy and rank order statistics were consistently obtained. Using the two-layer perceptron in a context-independent manner, a top-choice accuracy of 54% was achieved, which compares favorably with results reported in the literature. Context-dependent experiments reveal that heterogeneous sources of information can be integrated to improve recognition performance. The top-choice accuracy of 67% is comparable to the average agreement among listeners. Finally, it was found that the rate of improvement on recognition accuracy may be used as a terminating criterion for training, and that reasonable performance can be achieved using as few as 800 training tokens. [Work supported by DARPA.]