Building a sense-distinguished multilingual lexicon from monolingual corpora and bilingual lexicons.
Marcus Sammer, Stephen Soderland · 2007
Both lexical translation and knowledge-based translation systems require sense-distinguished translation lexicons, yet such lexicons are expensive to create manually. However, the abundance of untagged monolingual corpora and the availability of bilingual, machine-readable dictionaries (MRDs) suggest an opportunity. Our PanLexicon system takes advantage of these resources to automatically construct a sense-distinguished multilingual lexicon. The challenge for PanLexicon is that free, bilingual MRDs do not make sense distinctions, and often have spotty coverage. PanLexicon uses word contexts from monolingual corpora to guide it in finding translation sets – sets of words that share the same word sense across multiple languages. By maintaining word sense distinctions, PanLexicon finds translations between language pairs that are not supported by any of its bilingual source dictionaries. PanLexicon runs in time linear in the size of its input, and thus scales readily to large numbers of languages. We built a prototype of PanLexicon with inputs from Spanish-English and Chinese-English dictionaries. Our initial experimental results show that PanLexicon is able to find high-quality translation sets despite the limitations of its inputs. 1.