Complement the comparable corpus obtained from websites

Zhou Youliang, Gong Zhengxian, Zhou Guodong · 2010

This paper proposes a method to automatically extract high quality phrase translation tuples from web corpora, and discuss the automatic way to complement the lost part of the bilingual corpora for the first time. It analyzes the features of bilingual translation pairs in web pages, and then a statistical discriminative model combined with multiple features is used to extract translation pairs. Experimental results show that after our experiment, the corpus is aligned well enough for related research.

Read the paper · More papers on PaperTik