F0 contour parametric modeling using multivariate adaptive regression splines for arabic text-to-speech synthesis
Zied Mnasri, Fatouma Boukadida, Noureddine Ellouze · 2011
Arabic text-to-speech synthesis needs to be developed, in order to be integrated to many IT applications, like email and SMS reading, automatic information delivery and helping disabled people to use such sophisicated services. However, a standalone text-to-speech system needs automatic generation of prosody, including F0contour prediction. Thus, F0contour is linked to the text data via the Fujisaki model, which divides F0contour into phrase and accents components. Furthermore, the parametric structure of Fujisaki model reduces the problem into the estimation of parameters. Hence, regression techniques, such as MARS, are useful to map the text-retrieved features to the speech-signal-extracted parameters. Then, the overall F0contour is reconstructed and compared to the original one, to validate the model.