Sub-Sentence Alignment of Chinese-English Law Literature Based on Statistical Approach

Lifeng Xue · Journal of Northeastern University · 2003

Sentence alignment based on statistical approach is the choice of alignment with maximum probability from all candidates according to the length of bilingual sentences. ChineseEnglish law literature is translated literally, so it is suitable to be aligned with statistical approach. But the method used to compute the parameters in processing IndoEuropean languages cannot be applied to ChineseEnglish corpora. Two parameter computation methods for aligning ChineseEnglish corpora were presented. The method make the evaluation function satisfy the standard normal distribution. One method to get the parameter s2 is to compute slope of the line generated by linear regression analysis to all point (l1,(l2-cl1)2) in the training corpora. The other is to compute the variance. Test results show that the precision rate and recall rate of alignment are 98.8% and 99.2 % respectively.

Read the paper · More papers on PaperTik