A Statistical Word-Level Translation Model for Comparable Corpora

Mona Diab, Steve Finch · 2000

In this paper, we present a model of statistical word-level mapping for comparable corpora. The approach is based on the assumption that if two terms have close distributional profiles, their corresponding translations' distributional profiles should be close in a comparable corpus. The proposed model is described. A preliminary investigation on intralanguage comparable corpora is laid out. The preliminary results are >92% accurate, suggesting the feasibility of the model. The model needs to undergo some improvements and should be tested cross linguistically before assessing its significance. Keywords Word-level mapping, comparable, parallel, Spearman correlation, contingency, gradient descent 1. Introduction The natural language processing community is in constant need of readily available resources such as corpora, thesauri, bilingual and multilingual lexicons and dictionaries. The acquisition of such resources has proven to be challenging so far, requiring an immense overhead in...

Read the paper · More papers on PaperTik