Computational models of speech sound generation
James L. Flanagan · The Journal of the Acoustical Society of America · 1978
A comprehensive characterization of the acoustics of speech sound generation is central to the techniques of analysis/synthesis telephony. The classical model of linearly separable excitation source and resonant acoustic system has, for many years, served vocoder technology and speech-synthesis efforts. Now, analytical understanding of vocal-system acoustics is more extensive, is better quantified, and can be more readily represented in digital form. Improvements may ensue, therefore, from adopting a more sophisticated—yet computationally tractable—model of the speech signal. We review here analytical formulations that aim to describe the dominant acoustic effects in speech production. In particular we include models to represent the self-oscillating properties of the vocal cords, the acoustic interaction between cords and tract, noise generation from turbulent air flow, sound radiation from the mouth, nose and yielding tract walls, and the dynamics of vocal-tract shape. Finally, we offer preliminary examples of how the composite system, formulated for computer simulation, can be utilized for speech analysis and for speech synthesis.