Neural networks for text-to-speech phoneme recognition
Mark J. Embrechts, Fabio A. Arciniegas · 2002
Presents two different artificial neural network (ANN) approaches for phoneme recognition for text-to-speech applications: staged backpropagation neural networks and self-organizing maps. Several current commercial approaches rely on an exhaustive dictionary approach for text-to-phoneme conversion. Applying neural networks to phoneme mapping for text-to-speech conversion creates a fast distributed recognition engine. This engine not only supports the mapping of missing words in the database, but it can also mitigate contradictions related to different pronunciations for the same word. The ANNs presented in this work were trained based on the 2,000 most common words in American English. Performance metrics for the 5,000, 7,000 and 10,000 most common words in English were also estimated to test the robustness of these neural networks.