Emphasized Accent Phrase Prediction from Text for Advertisement Text-To-Speech Synthesis
Hideharu Nakajima, Hideyuki Mizuno, Sumitaka Sakauchi · Institutional Repositories DataBase (IRDB) · 2014
Realizing expressive text-to-speech synthesis needs both text processing and the rendering of natural expressive speech. This paper fo-cuses on the former as a front-end task in the production of synthetic speech, and in-vestigates a novel method for predicting em-phasized accent phrases from advertisement text information. For this purpose, we exam-ine features that can be accurately extracted by text processing based on current Text-to-speech synthesis technologies. Among fea-tures, the word surface string of the main con-tent and function words and the part-of-speech of main function words in an accent phrase are found to have higher potential on predict-ing whether the accent phrase should be em-phasized or not through the calculation of mu-tual information between emphasis label and features of Japanese advertisement sentences. Experiments confirm that emphasized accent phrase prediction using support vector ma-chine (SVM) offers encouraging accuracies for the system which requires emphasized ac-cent phrase locations as context information to improve speech synthesis qualities. 1