Voice quality conversion in TD-PSOLA speech synthesis
Xuejing Sun · 2002
The capability of producing different voice qualities is highly desirable in modern speech synthesis systems. Diphone based synthesizers using TD-PSOLA can generate high quality synthetic speech. However, one of the drawbacks of such systems in comparison to the formant synthesizer or the LPC synthesizer is its inflexibility in voice quality conversion (VQC). In this paper, the author presents a VQC method for the TD-PSOLA synthesis system. For vocal fry, the ST-signals are multiplied by a Kaiser window with alternate magnitude; for breathy voice, the ST-signals are first convolved with a one-pole filter, and then combined with shaped noise signals, and finally multiplied by a Hanning window. All the windowed ST-signals are then overlap-added as in standard TD-PSOLA. The perceptual evaluation test shows that this method can generate the desired voice quality successfully.