Adjacency Analysis for Designing Unit Selection Speech Model on Micro-Prosodic Level
Nur Hana Samsudin, Tang Enya Kong, Kim Chuah · 2006
One of the speech synthesizer problems is the unnaturalness and unintelligible production of speech. It is believed that wave modification could also contribute to the distortion of synthesized speech. The objective of a unit selection speech corpus is to provide a few possible instance of unit in order to produce synthesized speech as close to human speech production without the need (or slight need) to perform wave modification. Different from a standard prerecorded speech database, unit selection allow multiple instances for same unit to be presented in a corpus. These instances are differentiated by a unique combination of acoustic and linguistic value. These acoustic and linguistic values form a set of speech corpus parameter. To minimize distortion in concatenative speech synthesizer, parameters of selected segments need to reflect the quality of natural speech. The selection process depends on the priority level of each parameter in the corpus. These parameters design unit selection model. One of the proposed parameter is adjacent phoneme of the target unit. This paper will describe the analysis on micro-prosody behaviour particularly in pitch. We will also see how adjacency influenced the changes in micro-prosodic features and how the physics of speech contribute to the changes in micro-prosodic aspect.