Clustered language models based on regular expressions for SMT
Saša Hasan, Hermann Ney · RWTH Publications (RWTH Aachen) · 2005
In this paper, we present a language model based on clusters obtained by applying regular expressions to the training data and, thus, discriminating several different sentence types as, e.g.interrogatives, imperatives or enumerations.The main motivation lies in the observation that different sentence types also underlie a different syntactic structure, and thus yield a varying distribution of n-grams reflecting their word order.We show that this assumption is valid by applying the models to English-Spanish bilingual corpora and obtaining good perplexity reductions of approximately 25%.In addition, we perform an n-best rescoring experiment and show a relative improvement of 4-5% in word error rate.The models can be easily adapted to other translation tasks and do not need complicated training methods, thus being a valuable alternative for on-demand rescoring of sentence hypotheses such as they occur in the CAT framework.