Exploiting Phrasal Lexica and Additional Morpho-syntactic Language Resources for Statistical Machine Translation with Scarce Training Data

Maja Popović, Hermann Ney · RWTH Publications (RWTH Aachen) · 2005

Abstract. In this work, the use of a phrasal lexicon for statistical machine translation is proposed, and the relation between data acquisition costs and translation quality for different types and sizes of language resources has been analyzed. The language pairs are Spanish-English and Catalan-English, and the translation is performed in all directions. The phrasal lexicon is used to increase as well as to replace the original training corpus. The augmentation of the phrasal lexicon with the help of additional monolingual language resources containing morpho-syntactic information has been investigated for the translation with scarce training material. Using the augmented phrasal lexicon as additional training data, a reasonable translation quality can be achieved with only 1000 sentence pairs from the desired domain. 1 Introduction and Related Work The goal of statistical machine translation (SMT) is to translate an input word sequence f J 1 = f1... fj... fJ into a target word sequence eI 1 = e1... ei... eI by maximising the probability

Read the paper · More papers on PaperTik