Statistical Machine Translation - The Problem of Unknown Words

Vasques Filipe da Silva · 2012

Developing systems that translate words as accurately as humans is not an easy task. Statistical Machine Translation systems base themselves on training data. So, when translating a document, some of the words in that document might not have been encountered in the training phase and, thus, the system does not know how to translate these words. The objective of this work is to develop a system that finds possible translations of these unknown words. Since words from closely related languages have common ancestors, many of these words will end up having similarities that can help us in discovering if they are possible translations of each other or not, these words are named cognate words. Therefore, we explore orthographic similarities between words to find translations. We also make use of Logical Analogy, by attempting to infer the translation of the unknown word by looking at the translations of words related to it. Our final method tested uses the context in which each word is inserted to calculate how similar two words are. By merging these systems, we try to maximize the number of unknown words translated. Our approach is tested in the translation of corpora from Portuguese to English (and viceversa).

Read the paper · More papers on PaperTik