Synthesis by rule of a female voice

Dennis H. Klatt · The Journal of the Acoustical Society of America · 1981

Most speech synthesis-by-rule systems mimic the speech of a male talker. The reason is simple—the speech analysis tools that are available measure formant frequencies with greater accuracy when fundamental frequency is low, as in a male voice, and formant frequency data is essential when driving a formant synthesizer by rules. The research to be described attempts to convert an existing rule system (KlaTTalK) to a female voice quality through a judicious combination of theory and measurements on a particular female voice. The following steps were taken to this end: Fundamental frequency was scaled by a simple multiplicative factor of 1.8. Formant frequencies for vowels were scaled initially by “k factors” published by Fant to take into account differences in average formant sensitivity to changes in the length of the oral and pharyngeal vocal tract between males and females. The scaled formant values were then modified in some cases to better match observed data from the selected speaker. K factors were also extrapolated to apply to consonant articulations. Then formant bandwidth targets and parallel-branch amplitude targets were adjusted so as to match relative formant peak heights in selected consonant—vowel syllables. A demonstration tape will be played of the resulting synthetic voice.

Read the paper · More papers on PaperTik