Two Approaches to Matching in Example-Based Machine Translation

Sergei Nirenburg, Constantine Domashnev · 2005

This paper describes two approaches to matching input strings with strings from a translation archive in the example-based machine translation paradigm- the more canonical "chunking + matching + recombination " method and an alternative method of matching at the level of complete sentences. The latter produces less exact matches while the former suffers from (often serious) translation quality lapses at the boundaries of recombined chunks. A set of text matching criteria was selected to reflect the trade-off between utility and computational price of each criterion. A metric for comparing text passages was devised and calibrated with the help of a specially constructed diagnostic example set. A partitioning algorithm was developed for finding an optimum "cover " of an input string by a set of best-matching shorter chunks. The results were evaluated in a monolingual setting using an existing MT post-editing tool: the distance between the input and its best match in the archive was calculated in terms of the number of keystrokes necessary to reduce the latter to the former. As a result, the metric was adjusted and an experiment was run to test the two EBMT methods, both on the training corpus and on the working corpus (or "archive") of some 6,500 sentences. 47 The growth rate of theoretical studies of language structure and use stubbornly remains higher than the improvement

Read the paper · More papers on PaperTik