Context-dependent additive log F0 model for HMM-based speech synthesis

Heiga Zen, Norbert Braunschweiler · 2009

Abstract Thispaperproposesacontext-dependentadditiveacousticmod-elling technique and its application to logarithmic fundamentalfrequency (logF 0 ) modelling for HMM-based speech synthe-sis. Intheproposedtechnique,meanvectorsofstate-outputdis-tributions are composed as the weighted sum of decision tree-clustered context-dependent bias terms. Its model parametersand decision trees are estimated and built based on the maxi-mumlikelihood(ML)criterion. Theproposedtechniquehasthepotential to capture the additive structure of logF 0 contours. Apreliminary experiment using a small database showed that theproposed technique yielded encouraging results. Index Terms : speech synthesis, HMMs, logF 0 modelling 1. Introduction Hidden Markov model (HMM)-based speech synthesis [1] hasgrowninpopularityinrecentyears. Inthisframework,thespec-trum, excitation, and durations of speech are modelled simul-taneously in a unified framework of HMMs. For a given textto be synthesized, speech parameter trajectories that maximisetheir output probabilities are generated from estimated HMMsunder constraints between static and dynamic features [2]. Typ-ical instances of this framework use mel-cepstral coefficientsor line spectral pairs for their spectral parameters and logF

Read the paper · More papers on PaperTik