Readability Consideration in Speech Synthesis Recording Script Selection.

Minghui Dong, Ling Cen, Paul Chan, Haizhou Li · 2009

Designing text scripts that cover enough phonetic units and prosodic phenomena is very important when recording speech database for corpus based speech synthesis. When designing recording scripts for speech synthesis databases, a lot of effort is often placed on how to achieve maximal coverage of phonetic units in minimal speech recording. However, when we try to select sentences that have optimal coverage of the speech phenomena, some sentences with difficult words or incorrect grammar are often selected. It is difficult for speakers to read these sentences correctly and naturally at the same time. In order to address the problem in building speech database, we propose a selection process to create easy-to-read text scripts for recording in this paper. In this work, we will consider how to build a candidate set that is easy to read so that the speaker can utter it in the most natural way. We will calculate the statistics of the English text by analyzing the English Gigaword corpus, and filer out the sentences containing infrequent words and bigrams. The experiment shows that the selected scripts have good unit coverage of the language and good readability.

Read the paper · More papers on PaperTik