Bilingual word alignment of multi-strategy
Cai Dong-feng · Jisuanji gongcheng yu sheji · 2009
Aligning bilingual corpus at the word level is very important to statistical machine translation (SMT). The diversity and feasibility of morphology, semantics and syntax, with out-of-vocabulary words and segmentation error directly or indirectly affect the word alignment. An efficient multi-strategy alignment algorithm is presented, by combining the lexical information, GIZA++ results and HowNet. A set form operation is used to guide the disambiguation process of word alignment, according to the analysis of the bilingual corpus and the alignment result. The experiments show that F-score is 85.07% and increased by 10% over optimized IBM model, and alignment error ratio is decreased by 10%. The strategy complements the advantages of those algorithms according to the reliability and consistence of them.