Rapid Ramp-up for Statistical Machine Translation: Minimal Training for Maximal Coverage

Hemali Majithia, Philip Rennart, Evelyne Tzoukermann · 2005

This paper investigates optimal ways to get maximal coverage from minimal input training corpus. In effect, it seems antagonistic to think of minimal input training with a statistical machine translation system. Since statistics work well with repetition and thus capture well highly occurring words, one challenge has been to figure out the optimal number of “new ” words that the system needs to be appropriately trained. Additionally, the goal is to minimize the human translation time for training a new language. In order to account for rapid ramp-up translation, we ran several experiments to figure out the minimal amount of data to obtain optimal translation results. 1

Read the paper · More papers on PaperTik