Digital speech synthesis: Tutorial

Horabail S. Venkatagiri, Tenkasi V. Ramabadran · Augmentative and Alternative Communication · 1995

Digital speech synthesis typically involves conversion of text typed at a keyboard or input through an alternative method into digitally coded speech. Synthesis of speech using waveform coding is well understood, requires simple hardware and software, and results in high-quality, natural-sounding speech. However, waveform coding is memory intensive, which limits its use to restricted text-to-speech (TTS) output and digital recording and playback of several seconds worth of speech. Two parametric coding techniques, formant coding and linear predictive coding, are suitable for unrestricted TTS synthesis because of their relatively small memory overhead. Parametrically synthesized speech varies widely in quality and intelligibility from poor to very good, depending on the size of the linguistic units used for synthesis (e.g., word concatenation produces better results than phoneme concatenation) and on algorithmic sophistication, especially that used for handling text to phonetic code conversion and transitions across segment boundaries. While the best of the current speech synthesis algorithms produce highly intelligible speech, natural-sounding, unrestricted TTS has proved elusive. The present paper presents a broad overview of the components and processes of digital speech synthesis and discusses the advantages and disadvantages of different coding techniques used in speech synthesis.

Read the paper · More papers on PaperTik