Using estimated formants tracks for formants smoothing in text to speech (TTS) synthesis
Phuay Hui Low, C. -H. Ho, S. Yaseghi · 2004
Spectral or formant discontinuities across successive speech segments are legacies of concatenative TTS synthesisers. In this paper, a pole analysis procedure is used to estimate the formant frequency, bandwidths, spectrum shape and dynamics. The obtained formant tracks are then used for formant smoothing purposes in TTS synthesis. This paper explores three methods of spectral and formants smoothing. The first method achieves spectral smoothing at segment boundaries by interpolating the LP autocorrelation vectors. The second and third formant smoothing methods involves direct modification of the formant frequencies. In the second method, smoothing is achieved by substituting the formants of the synthesised speech with that of the natural speech. Finally, the last method achieves formant smoothing by moving the formants of successive segments closer to an average value.