A Phrase Table Filtering Model Based on Binary Classification for Uyghur-Chinese Machine Translation
Chenggang Mi, Yating Yang, Xi Zhou, Lei Wang, Xiao Li, Eziz Tursun · 2014
Abstract—In statistical machine translation, large amount of unreasonable phrase pairs in a phrase table can affect the decoding efficiency and the overall translation performance, especially in Uyghur-Chinese machine translation. In this paper, we present a novel phrase table filtering model based on binary classification, which consider differences between Uyghur and Chinese, and draw lessons from binary classification in machine learning. In our model, four features are considered: 1) Difference in length between source and target phrase; 2) Proportion of translated words in phrase pairs; 3) Proportion of symbol words; 4) Average number of co-occurrence words in training corpus. We use this model to generate a filtered phrase table. Experimental results show that this new filtering model can improve the performance and efficiency of our current Uygur-Chinese machine translation system. Index Terms—Uyghur-Chinese machine translation; Phrase table filtering; Binary classification