Source Error-Projection for Sample Selection in Phrase-Based SMT for Resource-Poor Languages
Sankaranarayanan Ananthakrishnan, Shiv Naga Prasad Vitaladevuni, Rohit Prasad, Prem Natarajan · 2011
The unavailability of parallel training cor-pora in resource-poor languages is a ma-jor bottleneck in cost-effective and rapid deployment of statistical machine transla-tion (SMT) technology. This has spurred significant interest in active learning for SMT to select the most informative sam-ples from a large candidate pool. This is especially challenging when irrelevant outliers dominate the pool. We propose two supervised sample selection methods, viz. greedy selection and integer lin-ear programming (ILP), based on a novel measure of benefit derived from error anal-ysis. These methods support the selec-tion of diverse and high-impact, yet rel-evant batches of source sentences. Com-parative experiments on multiple test sets across two resource-poor language pairs (English-Pashto and English-Dari) reveal that the proposed approaches achieve BLEU scores comparable to the full sys-tem using a very small fraction of all avail-able training data (ca. 6 % for E-P and 13% for E-D). We further demonstrate that the ILP method supports global constraints of significant practical value. 1