A new algorithm for speech synthesis based on vocal tract modeling
Qiguang Lin, Gunnar Fant · The Journal of the Acoustical Society of America · 1990
A new algorithm for articulatory speech synthesis is described in this paper. The algorithm constitutes two main parts: a detailed frequency domain modeling of the vocal tract and a data transform from the frequency domain to the time domain. A computer model was developed to simulate the vocal tract acoustics. The model incorporates all known important components of the vocal system and computes the transfer function between the lip/nostril output and the acoustic source. By decomposing the obtained transfer function into its numerator and denominator, frequencies and bandwidths of resonances (and of antiresonances, if any) can be determined. The transfer function can next be written as a partial fraction expansion series in terms of calculated residues at the poles and can be approximated by retaining the first few terms in the series. Usually, each of these terms is a second-order module and corresponds to an elementary formant resonance. Formants are thus connected in parallel. The time domain output is obtained by the inverse Laplace transform. Compared with other synthesis methods, this vocal-tract-oriented synthesis strategy has a number of advantages. For instance, the frequency dependency of loss elements of the vocal tract is preserved and accurate frequency responses can be reproduced. It is also computationally efficient relative to a direct convolution method. Examples of spectrum matching are presented to discuss the properties of the proposed algorithm. The algorithm has been incorporated in an articulatory-based speech synthesis system currently under development at KTH. [Work supported by grants form the Swedish Board for Technical Development (STU).]