Automatic labeling schemes for concatenative speech synthesis
Juraj Kačur, Jozef Cepko, Andrej Palenik · International Symposium ELMAR · 2008
This article discusses problems and solutions related to the labeling of the speech which is further used in speech synthesis. Although there are several synthesis methods, here we put focus on the concatenative speech synthesis, which is especially sensitive to labeling errors. As it uses huge amount of data it must be processed automatically. It is well accepted that the most precise automatic labeling methods are based on the automatic speech recognition systems and the most successful ones are utilizing HMM technology which connects more advantages, however there are other wide-spread methods like DTW and its various combinations. In the following paper HMM algorithm using various models was trained, tested and evaluated on couple of Slovak speech databases, which had either non-specific content or constructed to cover specific tasks such as weather forecast and railway information system. The results showed that the tested complexity of HMM models do not influence the accuracy of the labeled phoneme borders considerably. However, the comparison with a sample of hand-labeled data showed that the automatic HMM methods must be manually revised to be used for high quality concatenative synthesis.