Text‐to‐Speech Synthesis Development

Caroline G. Henton · The Encyclopedia of Applied Linguistics · 2012

Abstract Speech synthesis, as defined inThe International Encyclopaedia of Linguistics, “is the automatic generation of speech using linguistically salient acoustic or articulatory properties, or spoken units that are selected and controlled using computational commands” (Clark & Henton, 2003). Machines that synthesized speech (synthesizers) were unveiled first in 1939 at the New York World's Fair. In the following 50 years, early vocoder synthesis was improved upon by parametric synthesis and linear predictive coefficient (LPC) synthesis. Parametric (or formant) synthesizers get control information either by analyzing relevant acoustic parameters in speech, or by rules operating on a character string (i.e., text). LPC synthesis uses the representation of the speech signal as a set of coefficients that try to predict the signal from past values in the time domain; LPC accounts for resonance and pitch movements well, but fails in natural voice quality because of the invariant nature of glottal pulses. For a history of the development of, and progress in, speech synthesis, see Klatt (1987), van Santen, Sproat, Olive, and Hirschberg (1997), and Clark and Henton (2003).

Read the paper · More papers on PaperTik