High quality text-to-speech synthesis: a comparison of four candidate algorithms
Thierry Dutoit · 2002
We investigate the use of four candidate speech models in the context of high quality text-to-speech systems (HQ-TTS), address problems typically encountered by their prosody matching and segment concatenation modules, and compare their performances regarding: the segment database compression ratio they allow, the computational load of the related synthesis algorithms, as well as their intelligibility and subjective segmental quality. The models addressed are: the classical auto-regressive (LPC) one, the hybrid harmonic/stochastic (H/S) model proposed by Griffin and Lim (1988) and by Abrantes, Marques and Transcoso (1991), the 'null' model, as implemented by the time-domain pitch-synchronous overlap-add (TD-PSOLA) synthesis algorithm, and the multi-band re-synthesis pitch-synchronous overlap-add (MBR-PSOLA) model.>