Choosing the best machine translation system to translate a sentence by using only source-language information
Felipe Sánchez-Martínez · Repositorio Institucional de la Universidad de Alicante (Universidad de Alicante) · 2011
This paper describes a novel approach aimed to identify a priori which subset of machine translation (MT) systems among a known set will produce the most reliable translations for a given source-language (SL) sentence. We aim to select this subset of MT systems by using only information extracted from the SL sentence to be translated, and without access to the inner workings of the MT systems being used. A system able to select in advance, without translating, that subset of MT systems will allow multi-engine MT systems to save computing resources and focus on the combination of the output of the best MT systems. The selection of the best MT systems is done by extracting a set of features from each SL sentence and then using maximum entropy classifiers trained over a set of parallel sentences. Preliminary experiments on two European language pairs show a small, non-statistical significant improvement.