An n-gram Approach to Exploiting a Monolingual Corpus for Machine Translation
Toni Badía, Gemma Boleda, Maite Melero, Antoni Oliver · O2 - Repositori Institucional (Universitat Oberta de Catalunya) · 2005
In this paper we present an approach to Statistical Machine Translation that uses a bilingual dictionary and a target language model based on n-grams extracted from a monolingual corpus. This approach is still in an experimental stage and is being developed in the context of Metis-II, a UE project that aims at constructing free text translations by retrieving the basic stock for translations from large monolingual corpora. The architecture described in this paper is being applied to translation from Spanish to English and is designed so as to depend as little as possible on complex linguistic processing tools. The only required tools are a POS tagger and lemmatizer for the source language an another for the target language.