Speech production using neural networks and concatenated phonetics

William Bruce Hudson, J. Eldon Steelman · 1990

The production of synthetic speech using an inventory of 835 diphone speech segments is presented. Diphone and phoneme generated speech are evaluated using the hardware designed for this study. It has been found that the use of concatenated phonemes (diphones) produces more natural and understandable speech than speech constructed using the 45 phoneme inventory normally employed in synthetic speech production. The use of artificial neural networks to perform text to speech production was evaluated. While studies presented demonstrate that the learning method used to train the network can impact on how quickly the network learns associations, it was also found that excessive learning rates can impact on system stability. The present platform for text to phoneme mapping, the IBM PC has been found not to be able to in a real time sense, accommodate the neural network for text to speech translation.

Read the paper · More papers on PaperTik