TakeTwo: A Word Aligner based on Self Learning
Jim Chang, Jian-Cheng Wu, Jason S. Chang · Institutional Repositories DataBase (IRDB) · 2014
State of the art statistical machine translation systems are typically trained by symmetrizing word alignments in two translation directions. We introduce a new method that improves word alignment results, based on self learn- ing using the initial symmetrized word align- ments results. The method involves align- ing words and symmetrizing alignments, gen- erating labeled training data, and construct a classifier for predicting word-translation rela- tion in another alignment round. In the first alignment round, we use the original grow- diag-final-and procedure, while in the second round, we use the classifier and a modified GDFA procedure to validate and fill in align- ment links. We present a prototype system, TakeTwo, which applies the method to im- prove on GDFA. Preliminary experiments and evaluation on a hand-annotated dataset show that the method significantly increases the pre- cision rate by a wide margin (+16%) with comparable recall rate (-3%).