Segmenting unrestricted Chinese text into prosodic words instead of lexical words
Yao Qian, Min Chu, Hu Peng · 2002
This paper stresses the importance of converting a string of lexical words to that of prosodic words in text-to-speech (TTS) systems by presenting the surface differences and perceptual differences between them. A statistical rule based method and a classification and regression tree (CART) based method are proposed as solutions. Though ComplicatedSet based CART method performs the best, the achievement is obtained at the cost of heavy computation workloads needed by a parser. Statistical rule based method results in higher recall but lower precision, comparing to SimpleSet CART method. It is very difficult to tell which is better, since we don't know which affects naturalness more, precision or recall. Both of them require only lexicon word segmentation and part of speech (POS) tagging in the preprocessing stage, and are easily realized in TTS systems. Results of the preference test discloses that significant improvements on naturalness are perceived when lexical word strings are converted into prosodic word strings by our approach.