Chinese-Uighur Sentence Alignment Based on Hybrid Strategy with Mistake Spread Suppression
Shengwei Tian, Turgun Ibrahim, Hasan Umal, Long Yu · 2009
This paper proposes a hybrid algorithm based on mistake spread suppression to align Chinese-Uighur sentences. Aiming at the shortcoming of mistake spread in alignment algorithm based on length, this paper presents a new kind of suppression strategy for mistake spread. This strategy omits Chinese segmentation and processing for post tagging. By using characteristics of punctuation, sentence length and Chinese-Uighur correspondence information,the anchor points with 1:1 pattern sentence pairs are identified to suppress mistakes spread. Among anchor points, a hybrid strategy based on both length and punctuation is used to align sentences. Experimental results verified the high precision of identifying anchor points and the effective restraint of the spread of alignment mistakes; Hybrid alignment algorithm avoids the weakness of high time complexity alignment algorithms based on word. In addition, its performance is improved more compare with traditional alignment algorithms, and alignment mistake ratio is reduced from 4.8% to 2.3%.