Bilingual chunk alignment in statistical machine translation

Zhou Yu, Chengqing Zong, Bo Hao Xu · 2005

In this paper a new algorithm called multilayer filtering (MLF) is proposed for extracting bilingual alignment chunks automatically from a Chinese-English parallel corpus. Multiple layers are used to extract bilingual chunks according to different features of chunks in the bilingual corpus. And the alignment chunks are one-to-one corresponding with each other. The chunking and alignment algorithm doesn't rely on the information from tagging, parsing, syntax analyzing or segmenting for Chinese corpus as most conventional algorithms do. Preliminary experimental results show that the algorithm achieves a good performance in chunking and alignment. Moreover, the translations generated by this algorithm are much better than the results generated by the baseline (word-based statistical machine translation).

Read the paper · More papers on PaperTik