Concatenation cost calculation and optimisation for unit selection in TTS
2002
This paper presents a concatenation cost used during the selection of acoustic units in speech synthesis. The concatenation cost is defined as a linear function of weighted sub-costs, each addressing a particular phonetic or prosodic feature. The different sub-costs composing the concatenation cost are defined, and their weights are optimised by a multiple linear regression as a function of an acoustic measure of concatenation quality. A perceptual evaluation of the synthesised speech is done with the selection obtained using optimised and hand-tuned weights and with the current laboratory's unit selection method.