Sub-Sentence Alignment of Chinese-English Law Literature Based on Statistical Approach
Lifeng Xue · Journal of Northeastern University · 2003
Sentence alignment based on statistical approach is the choice of alignment with maximum probability from all candidates according to the length of bilingual sentences. ChineseEnglish law literature is translated literally, so it is suitable to be aligned with statistical approach. But the method used to compute the parameters in processing IndoEuropean languages cannot be applied to ChineseEnglish corpora. Two parameter computation methods for aligning ChineseEnglish corpora were presented. The method make the evaluation function satisfy the standard normal distribution. One method to get the parameter s2 is to compute slope of the line generated by linear regression analysis to all point (l1,(l2-cl1)2) in the training corpora. The other is to compute the variance. Test results show that the precision rate and recall rate of alignment are 98.8% and 99.2 % respectively.