Using multiple linguistic features for Mandarin phrase break prediction in maximum-entropy classification framework

Yu Zheng, Gary Geunbae Lee, Byeongchang Kim · 2004

Abstract We model Mandarin phrase break prediction as a classification problem with three level prosodic structures and apply conditional maximum entropy classification to this problem. We acquire multiple levels of linguistic knowledge from an annotated corpus to become well-integrated features for maximum entropy framework. Five kinds of features were used to represent various linguistic constraints including POS tag features, lexical features, phonetic features, length features, and distance features. Experiment results show that our method performs better than the previous methods and the conditional maximum entropy (ME) model is very effective for data sparseness problem in Mandarin phrase break prediction. 1. Introduction Assigning the appropriate phrase breaks in text-to-speech systems is important for naturalness and intelligibility. Linguistic researchers have shown that the spoken language is structured as a hierarchy of prosodic units, including phonological phrase, intonation phrase, and utterance [1]. However, the text language is often structured by syntactic units, such as words and phrases, which are not equivalent to prosodic ones. But we suppose the syntactic information would provide important cues for prosodic phrase prediction. Many techniques have been introduced to predict phrase break, such as using Recurrent Neural Network (RNN)[2], Hidden Markov Model (HMM)[3], POS bi-gram and CART [4] and rule learning with C4.5 or TBL [5][6]. Min Chu and Yao Qian [4] proposed CART-based approach in four-class prosodic structure and their method shows high accuracy of 83%. Zhao and Tao with others [5][6] proposed automatic rule-learning approach with two typical rule-learning algorithms (C4.5 and TBL). They report a higher accuracy of 87.9% where they used POS features, lexical features and length features. They added chunking features and achieved even better accuracy of 90.0%, but they only used two-class prosodic structure to evaluate the accuracy. We treat entire phrase break prediction as a classification problem and apply conditional maximum entropy (ME) model. Various linguistic information is represented in the form of features, and five kinds of features were used in our system including POS tag features, lexical features, phonetic features, length features, and distance features. One serious problem of Mandarin phrase break prediction is that we usually do not have a large sized phrase break annotated corpus. So, we have to acquire multiple levels of linguistic information from only a small sized annotated corpus. Here, the data sparseness is a critical problem in Mandarin phrase break prediction, and we show that ME framework is very effective to capture useful features in the sparse training data environment. The remainder of this paper is organized as follows: Section 2 briefly introduces the ME framework with some justification. In section 3, we present our prosodic phrase framework and the five kinds of features used in the system. The effectiveness of our proposed method is verified by the experimental results given in section 4. Finally, conclusions are provided in section 5.

Read the paper · More papers on PaperTik