A neural network-controlled formant synthesizer with phoneme-dependent voicing
Michael S. Scordilis · 1991
Automatic formant synthesis has been an important method of text-to-speech conversion. However, two problems are currently contributing to its reduced quality: first, the nature of the excitation source during voicing, and second, the development of rules for synthesis from phonemic input. This dissertation addresses both problems and presents alternative solutions. For the voicing source, a production model is developed based on an equivalent to the human speech production system. The acoustical effects of the subglottal and supraglottal sections on the air volume velocity through the vibrating glottal folds are included in a new algorithm. Comparisons show that phoneme-dependent glottal pulses are superior to generic invariant pulses in producing the spectral characteristics of natural speech. A practical approximation suitable for real-time applications is also presented. For the rules to control a formant synthesizer, a new method is presented in which artificial neural network architectures were developed and trained on larynx-produced phonemic triplets and pairs from a naturally spoken set of words. The system was able to learn to control the cascade/parallel synthesizer model and produce new words that are intelligible and quite natural sounding. Because the performance of the proposed method is limited only by the size of the training corpus, if the required computing resources are available it can be trained extensively for very high performance, such as natural and unrestricted text-to-speech synthesis.