Extraction of Translation Equivalents from Parallel Corpora Using Sense-sensitive Contexts

Pablo Gamallo · 2005

Abstract. The paper proposes an unsupervised method to extract translation equivalents from parallel corpora. The strategy we use takes into account the context of words. Given a word of the source language and a particular context, we learn its word translation within an equivalent context. We first extract pairs of similar contexts and, then, we compare the similarity between words appearing in each pair. This allows us to use a very low threshold to identify correct translation equivalents. Moreover, as polysemic words tend to have different senses in different context pairs, we are able to associate several translation equivalents to the same polysemic word. The main contribution of this paper is precisely to learn the correct translation equivalent of a word in a specific context. On the other hand, we do not align texts by detecting sentences or other small linguistic units. We identify natural boundaries by detecting explicit parts or segments of the corpus. Most text corpora contain natural boundaries to explicitly separate basic parts such as chapters, articles, receipts, legal documents, letters, etc. We use these explicit and natural parts to align parallel corpora. To compute similarity within these large segments, we define a particular version of the Dice coefficient. 1.

Read the paper · More papers on PaperTik