MODELLING SPEECH TEMPORAL STRUCTURE FOR ESTONIAN TEXT-TO-SPEECH SYNTHESIS: FEATURE SELECTION

Meelis Mihkla · Trames Journal of the Humanities and Social Sciences · 2007

The article discusses the principles of selecting features for modelling the temporal structure of Estonian speech, using different types of read-out texts, with a view to text-tospeech synthesis (TTS).Feature selection is known to depend on certain general issues regulating speech temporal structure, as well as on some language specific aspects.The durational model of Estonian stands out for some foot-bound features (foot quantity degree, number of feet in the word) being included in the input.In addition to the traditional descriptors of sound context and hierarchical position the prediction of Estonian segmental durations requires information on some morphological, syntactic and lexical features of the word, such as word form, part of sentence, and part of speech.In the prediction of pauses in the speech flow the relevant features are: distance from sentence beginning and from the previous pause, the length and quantity degree of the preceding foot, and the occurrence of a punctuation mark or conjunction.Although expert opinions were used in feature selection, statistical methods should be applied to test the vector of optimal argument features.

Read the paper · More papers on PaperTik