Synonym Acquisition Using Bilingual Comparable Corpora

Daniel Andrade, Masaaki Tsuchida, Takashi Onishi, Kai Ishikawa · International Joint Conference on Natural Language Processing · 2013

Various successful methods for synonym acquisition are based on comparing context vectors acquired from a monolingual corpus. However, a domain-specific corpus might be limited in size and, as a consequence, a query term’s context vector can be sparse. Furthermore, even terms in a domain-specific corpus are sometimes ambiguous, which makes it desirable to be able to find the synonyms related to only one word sense. We introduce a new method for enriching a query term’s context vector by using the context vectors of a query term’s translations which are extracted from a comparable corpus. Our experimental evaluation shows, that the proposed method can improve synonym acquisition. Furthermore, by selecting appropriate translations, the user is able to prime the query term to one sense.

Read the paper · More papers on PaperTik