Speech synthesis system based on a variable decimation/interpolation factor
Francisco M. Gimenez de los Galanes, M. H. Savoji, José Manuel Pardo · 2002
In this paper we present a modification of the usual decimation-interpolation steps for resampling of speech signals which is especially adapted to arbitrary modification of fundamental frequency and duration of speech segments. The modification is intended to overcome the time and frequency domain limitation that such a resampling scheme imposes so it can be used in a speech synthesis system. The performance of this resampling method for prosody modification is better than the equivalent PSOLA (Pitch-Synchronous Overlap-Add) method when working at a sampling frequency of 8 to 10 kilohertz so the source spectrum of the voiced allophones can be said to be completely harmonical. An optimization of the proposed algorithm that allows a real time implementation is also discussed.