Speech synthesis by phonological structure matching

Paul R.P. Taylor, Alan W. Black · 1999

This paper presents a new technique for speech synthesis by unit selection. The technique works by specifying the synthesis target and the speech database as phonological trees, and using a selection algorithm which finds the largest parts of trees in the database which match parts of the target tree. The technique avoids many of the errors made by prosody generation modules by incorporating their operation in the selection implicitly. A technique for using signal processing only when it is needed most is also described. The technique produces better quality speech than previous approaches and is also significantly faster. 1. INTRODUCTION It is common in any overview of a speech synthesis system (e.g. [13], [8]) to see the system broken down into a number of components, which nearly always include things such as text normalisation, lexical lookup, intonation, duration, diphone concatenation and signal processing. A standard model of waveform generation over the last years has been fo...

Read the paper · More papers on PaperTik