On long sentence segmentation for C-E parallel patent corpora
Miao Xue-lei · Journal of Shenyang Aerospace University · 2011
Parallel corpora have proved extremely useful resources for the corpus-based machine translation.It is of great necessity to segment long sentences in Chinese-English parallel patent corpora to make it easy to use translation models in corpus-based machine translation systems.This paper puts forward a new statistical alignment method of long sentence segmentation,i.e.treating positions of punctuations as candidate split points,and then evaluates the segments by modified IBM model 1 which makes sentence segmentation more exactly.Experiments show that this method is effective to improve the statistical machine translation.