Fundamental frequency modeling for corpus-based speech synthesis based on a statistical learning technique

S. Sakai, James Glass · 2004

The paper proposes a novel two-layer approach to fundamental frequency modeling for concatenative speech synthesis based on a statistical learning technique called additive models. We define an additive F/sub 0/ contour model consisting of long-term (intonational phrase-level) component and short-term (accentual phrase-level) component, along with a least-squares error criterion that includes a regularization term. A backfitting algorithm, that is derived from this error criterion, estimates both components simultaneously by iteratively applying cubic spline smoothers. When this method is applied to a 7,000 utterance Japanese speech corpus, it achieves F/sub 0/ RMS errors of 28.9 and 29.8 Hz on the training and test data, respectively, with corresponding correlation coefficients of 0.81 and 0.77. The automatically determined intonational and accentual phrase components behave smoothly, systematically, and intuitively under a variety of prosodic conditions.

Read the paper · More papers on PaperTik