Producing a Cross-Language Dictionary using Statistical Machine Translation: A First Experiment with English and Indonesian

Joseph Cathcart, Robert Dale · 2001

Well-developed Statistical Machine Translation techniques now exist for carrying out word alignment in parallel corpora. A by-product of training data for this task is a set of translation probabilities for the correspondences between target and source tokens. In the literature, these techniques have relied on the use of bitexts of significant size; however, for many languages no such corpora exist. In this paper, we report on an experiment where a relatively small corpus was used to generate word correspondences, which were then evaluated against a hand-constructed bilingual lexicon. We present the suprisingly good results achieved, and discuss some possible improvements to the technique.

Read the paper · More papers on PaperTik