Towards Efficient Large-Scale Feature-Rich Statistical Machine Translation

Vladimir Eidelman, Ke Wu, Ferhan Türe, Philip Resnik, Jimmy Lin · 2013

We present the system we developed to provide efficient large-scale feature-rich discriminative training for machine translation. We describe how we integrate with MapReduce using Hadoop streaming to allow arbitrarily scaling the tuning set and utilizing a sparse feature set. We report our findings on German-English and Russian-English translation, and discuss benefits, as well as obstacles, to tuning on larger development sets drawn from the parallel training data. 1

Read the paper · More papers on PaperTik