Improving the Performance of GIZA++ Using Variational Bayes
Darcey Riley, Daniel Gildea · UR Research (University of Rochester) · 2010
Bayesian approaches have been shown to reduce the amount of overfitting that occurs when running the EM algorithm, by placing prior probabilities on the model parameters. We apply one such Bayesian technique, variational Bayes, to GIZA++, a widely-used piece of software that computes word alignments for statistical machine translation. We show that using variational Bayes improves the performance of GIZA++, as well as improving the overall performance of the Moses machine translation system in terms of BLEU score. GIZA++ (Och and Ney, 2003) is a program that trains the IBM Models (Brown et al., 1993) as well as a Hidden Markov Model (HMM) (Vogel et al., 1996), and uses these models to compute Viterbi alignments for statistical machine translation. While GIZA++ can be used on its own, it typically serves as the starting point for other machine translation systems, both phrase-based and syntactic.