Occurrence Based Statistics in Machine Translation

Vamshi Ambati, Alon Lavie · 2005

As MT approaches demand longer context for better translation quality, the limitations of current language modeling techniques become explicit. The computational inability to model the likelihood of longer ngrams and the likelihood of their usage in probabilistic manner, have prevented us from exploring long ngrams in MT. In this paper, we propose and investigate a new set of features called occurrence based statistics for machine translation that overcome these limitations. Occurrence based statistics are obtained by looking up the existence of sequence of lexical units (ngrams) in huge human created language repositories without actually dealing with their probabilities. We also experiment occurrence based statistics for certain syntactic units like the headword chains (nchains). Our experiments show that they correlate well with translation quality and are useful as discriminative features in classification too. 1

Read the paper · More papers on PaperTik