Integrating corpus-based and rule-based approaches in an open-source machine translation system
Felipe Sánchez-Martínez, Juan Antonio Pérez-Ortiz, Mikel L. Forcada · RUA, Repositorio Institucional de la Universidad de Alicante (Universidad de Alicante) · 2007
Most current taxonomies of machine translation (MT) systems start by contrasting rule-based (RB) systems with corpusbased (CB) ones. These two approaches are much more than theoretical boundaries since many working MT systems fall within one of them. However, hybrid MT systems integrating RB and CB approaches are receiving increasing attention. In this paper we show our current research on using CB methods to extend a MT system primarily designed following the RB approach. Specifically, the open-source MT system Apertium is being extended with a set of CB tools to be also released under an open-source license, therefore allowing third parties to freely use or modify them. We present CB extensions for Apertium allowing (a) to improve its part-of-speech tagger, (b) to automatically infer the set of transfer rules, and (c) to tackle the problem of the translation of polysemous words. A common feature of these CB methods is the use of unsupervised corpora in the target language of the MT system. The resulting hybrid system preserves most of the advantages of the RB approach while reducing the need for human intervention. 1