Development of a female synthetic voice using concatenative synthesis
Ann K. Syrdal · The Journal of the Acoustical Society of America · 1990
Although there is considerable demand for synthetic female voices in text-to-speech applications, current analysis and synthesis algorithms are more successful for male than for female speech. Difficulties include voice source characteristics and other acoustic differences between female and male speech. Telephone bandwidth constraints present an additional disadvantage for the female voice, because relatively more phonetically relevant high-frequency acoustic information is filtered out of the telephone signal for female speakers than for males. A synthetic female voice is currently being developed for use with an experimental AT&T concatenative text-to-speech system for which only male voices have been available previously [J.P. Olive, Proc. ICASSP, 568–570 (1977); J.P. Olive and M. Y. Liberman, J. Acoust. Soc. Am. Suppl. 1 78, S6 (1985)]. In this synthesis-by-rule system, segments used for synthesis are obtained from natural speech and concatenated to synthesize any arbitrary English utterance. Issues such as speaker selection, preparation and recording of speech materials, inventor? of concatenative units, and analysis and synthesis methods will be discussed. Diagnostic use of intelligibility test results and preliminary evaluations also will be described.