N-BEST REORDERING IN STATISTICAL MACHINE TRANSLATION
Francisco Casacuberta · 2015
As statistical machine translation (SMT) systems strive to improve the translation quality they are able to deliver, the word reordering problem is being unveiled as a major problem that must be addressed, whenever these systems are to be improved. While most works published focus their results in corpora involving English, Chinese and Arabic, such a translation problem can also be found within Spain itself: its origin being unknown, Basque presents a very peculiar word order, which is very different to most other european languages, and specially very different to Spanish word order. Because of this fact, SMT systems not including some sort of word reordering yield unsatisfactory results, involving serious training problems, when confronted with the Basque-Spanish task. Although some efforts have been made towards including word or phrase reordering in the decoding algorithms, these approaches usually imply a computational overhead that obliges the designers of such algorithms to assume sub-optimal restrictions, which often lead to a significant dimish in the translation quality. In this work, we present a reordering method based on the extraction and exploitation of monotized corpora, which prove to be specially useful for the language pairs presenting severe word reorderings. Our system has been tested on the Basque Tourist task, where very promissing results have been obtained. 1.