Exploration of acoustic correlates in speaker selection for concatenative synthesis
Ann K. Syrdal, Alistair D. Conkie, Yannis Stylianou · 1998
It is often difficult to determine the suitability of a speaker to serve as a model for concatenative text-to-speech synthesis. The perceived quality of a speaker's natural voice is not necessarily predictive of its (even relative) synthetic quality. The selection of female and male speakers on whom to base two synthetic voices for the new AT&T text-to-speech system was made empirically. Brief readings of identical text materials were recorded from pre-selected professional speakers (6 females, and 9 males). Small-scale TTS systems were constructed with a minimal diphone inventory, suitable for synthesizing a limited number of test sentences. Synthesized sentences, and their naturally spoken references, were presented to listeners in a formal listening evaluation. Listeners rated each test sentence independently on intelligibility, naturalness, and pleasantness. A variety of acoustic measurements of the speakers were made in order to determine which acoustic characteristics correlated ...