A genetic algorithm with look-ahead mechanism to estimate formant synthesizer input parameters

Jonathas Trindade, Fabíola Pantoja Oliveira Araújo, Aldebaro Barreto da Rocha Klautau Junior, Pedro Batista · 2013

There are several commercial text-to-speech (TTS) systems that generate speech signals that sound very natural. A distinct problem is utterance copy, which consists in taking speech as input (instead of text, as in TTS) and find the input parameters that would drive a speech synthesizer to generate speech that mimics the target speech with respect to contents and speaker identity. Utterance copy is a difficult task due to the need of adjusting several parameters and their nonlinear relation to the output. Genetic algorithms (GA) have been used in this task embedded in an analysis-by-synthesis loop, which requires solving several optimization processes, one for each short segment of speech. The contribution of this work is to present two new strategies, namely the voicing gene and the lookahead mechanism, which consistently improve the performance of the previous GA-based architecture. The results show that the proposed improvements reduced the mean squared error by more than 40%.

Read the paper · More papers on PaperTik