Moving Beyond Phrase Pairs: The Relevance of the Corpus in a SMT World

Aaron B. Phillips · 2010

Machine translation has advanced considerably in recent years, but primarily due to the availability of larger data sets. Translation of low-frequency phrases and resourcepoor languages is still a serious problem. In this work we explore a deeper integration of context, structure, and similarity within machine translation. Instead of modeling phrase pairs in abstract, we propose modeling each instance of a translation in the corpus. Unlike the traditional SMT approach that builds a mixture of independent, simple distributions for each phrase pair, our model is a mixture of translation instances. The significance lies in that we use a distance measure to assesses the relevance of each translation instance. It permits simple the integration of instance-specific features which we plan to exploit in three key directions. First, we will introduce non-local features that identify the relevant context of an instance in order to favor those that are most similar to the input. Second, we will mark-up the corpus with metadata from multiple external sources to sharpen the scoring of each translation instance and guide the overall translation process. Third, we will identify

Read the paper · More papers on PaperTik