Corpora in machine translation
Hanne Moa · 2005
In spite of this quote there are many machine translation systems in use today, and more are being made, as the need for translations is seemingly boundless. For instance the contracts, agreements, laws and parliamentary sessions of the EU need to be translated somehow, and even if machine translation as good as human translation is infeasible, as that is what Bar-Hillel was concerned about, even quickly made, partial and rudimentary translations can be of help to a translator, or to choose what texts need to be properly translated in the first place. Intriguingly, it turns out that corpora can help with the second claim of the quote, that “Computer understanding of text is too difficult.”. Corpora can provide some understanding of the world simply by being a source for deriving frequencies or other statistically significant phenomena like finding collocations and words that don’t follow the rules. In this paper I will zoom in from the general to the specific, going from the past up to today. Section 2 looks at early use of corpora and machine translation (hereafter MT), section 3 is about modern MT and its growing dependence on corpus linguistics and finally, section 4 is about a specific MT-project that is still under development and its use of corpora, namely LOGON (Lonning et al., 2004).