Issues in the development of the next generation of concatenative speech synthesis systems
Andrew Breen · 2000
While commercial organisations are as interested in the ability of text-to-speech (TTS) systems to perform well at text normalisation and pronunciation as they are in the overall quality of the synthetic voice, there is still a general opinion that TTS systems are only usable for limited applications. The vast majority of applications still require more natural speech and a greater variety of styles and emotion. Work on mark-up has attempted to address the limitations of a plain text interface, but has done little to solve the basic problems of how to maintain a high quality synthetic voice and apply different speaking styles. As a consequence this paper is concerned with one narrow aspect of TTS synthesis, that is the improvement of voice quality through the progressive improvement of existing approaches to unit selection.