Nonlinear time-scale modification of speech signal with varied segmental duration characteristics
Yoh’ichi Tohkura, Yoshinori Kitahara · The Journal of the Acoustical Society of America · 1988
Segmental duration of each phoneme changes depending upon the speaking rate. Generally, vowel parts are easier to be compressed or expanded than consonant parts are in fast or slow speech, respectively. Questions raised in this paper include how the speaking rate can be extracted from the speech signal without knowing the content (i.e., phonetic information) and what kind of time-scale modification can be chosen in order to control speaking rate. First, the segmental duration compressibility of the speech signal was defined by path slopes in DTW spectral matching when utterances with various kinds of speaking rates were matched to a reference utterance of a normal speaking rate. On the assumption that the compressibility is inversely proportional to segmental spectrum changes, the relationship between the compressibility and the average cepstral time difference Δcep [S. Furui, IEEE Trans. Acoust. Speech Signal Process. ASSP-34, 52–59 (1986)] was studied. The results showed that the Δcep is an efficient parameter to represent the compressibility. By using the Δcep as a control parameter of the speaking rate, nonlinear time-scale modification can be achieved without speech quality degradation.