A flexible and high quality articulatory speech synthesizer
Yu-Fu Hsieh, D. G. Childers · University of Florida Digital Collections (University of Florida) · 1994
The aim of this research was to develop one solution to the speech inverse filtering problem and to develop a flexible, high quality articulatory speech synthesis tool. The results of this study will be of interest to researchers in speech modeling, analysis, and synthesis. A software program called ARTM was implemented as an articulatory synthesis tool. One feature of this research tool is the simulated annealing optimization procedure that is used to optimize the vocal tract parameters to match a specified set of formant characteristics. Another aspect of this study is the derivation of a new form of the acoustic equations that include the subglottal system, the glottal impedance, the turbulence noise source, and the nasal tract with sinus cavities for the articulatory synthesizer. A flexible articulatory model was designed with special interfaces that provide for numerical specification of parameters as well as sliding bar capabilities that allow parameter adjustments. A transmission-line circuit model of the vocal system, which includes the vocal tract, the nasal tract with sinus cavities, the glottal impedance, the subglottal tract, the excitation source, and the turbulence noise source, was constructed. The acoustic equations of the vocal system were rederived for the proposed articulatory synthesizer. A digital time-domain approach was used to simulate the dynamic properties of the vocal system as well as to improve the quality of the synthesized speech. A new efficient analysis scheme, identifying the articulatory parameters from the acoustic speech waveforms, was induced. The algorithm is known as simulated annealing, which is constrained to avoid nonunique solutions and local minima problems. The constraints were determined by the articulatory-to-acoustic transformation function and the boundary conditions for the articulatory parameters. The cost function was defined as a percentage of the weighted least-absolute-value error distance between the first four formant frequencies of the articulatory model and the first four formant frequencies determined from speech analysis. A 1% error criterion was found to be both practical and achievable.