Mining new word translations from comparable corpora

Li‐Dong Shao, Hwee Tou Ng · 2004

New words such as names, technical terms, etc appear frequently.As such, the bilingual lexicon of a machine translation system has to be constantly updated with these new word translations.Comparable corpora such as news documents of the same period from different news agencies are readily available.In this paper, we present a new approach to mining new word translations from comparable corpora, by using context information to complement transliteration information.We evaluated our approach on six months of Chinese and English Gigaword corpora, with encouraging results.

Read the paper · More papers on PaperTik