Investigations on hierarchical phrase based machine translation

David Vilar Torres, Hermann Ney · 2011

In this thesis we investigate the hierarchical phrase-based approach to machine translation, with special attention to the search problem. This approach is nowadays one of the most widely applied for statistical machine translation, and thus a detailed study helps in advancing the state-of-the-art in the field. Two are the most widely used algorithms for translating using the hierarchical phrase-based approach: cube pruning and cube growing. For each of this algorithms we study their behaviour in terms of translation quality and computational requirements (speed and memory usage), and propose novel extensions which improve the computational costs of the generation process. These extensions enable us to apply the hierarchical approach to wider domains and allow the use of larger sets of parallel corpora, which in turn improve translation quality. Furthermore, we design extensions of the hierarchical model that include linguistically motivated information into the translation process, comparing them with other approaches proposed by other research groups. By inspecting the behaviour of one of these methods, which includes additional information in the form of syntactic constituents, we propose a generalization that retains the structural properties of the model, but substitutes the syntactic information by information derived from automatic clustering techniques. This allows the use of this method for a broader spectrum of languages, where the necessary linguistic tools for the original method may not be available. An additional result of this thesis is the open source machine translation toolkit Jane, which was made available to the scientific community, free of charge for noncommercial purposes. The methods described in this thesis are all implemented in the toolkit, which provides a wider dissemination of the results, as well as allowing better replicability. Some practical implementation aspects are also discussed in this thesis. In the second part of the thesis, we turn our attention to the evaluation of machine translation output, focusing on three concrete subtopics of this broad area. First, we propose a novel method for performing human evaluation based on binary comparisons, which aims at speeding up the time-consuming process of evaluating machine translation output by human judges. Second, we present a framework for the classification of errors in machine-generated translations, which allows to detect the main problems of a translation system and focus research efforts. Lastly, we give evidence about the lack of correlation between the alignment error rate measure and the final translation quality, thus motivating a better inspection of the improvements

Read the paper · More papers on PaperTik