Integrating Models Derived from non-Parametric Bayesian Co-segmentation into a Statistical Machine Transliteration System

Andrew Finch, Paul R. Dixon, Eiichiro Sumita · 2011

The system presented in this paper is based upon a phrase-based statistical machine transliteration (SMT) framework. The SMT system’s log-linear model is aug-mented with a set of features specifically suited to the task of transliteration. In par-ticular our model utilizes a feature based on a joint source-channel model, and a fea-ture based on a maximum entropy model that predicts target grapheme sequences using the local context of graphemes and grapheme sequences in both source and target languages. The segmentation for our approach was performed using a non-parametric Bayesian co-segmentation model, and in this paper we present ex-periments comparing the effectiveness of this segmentation relative to the publicly available state-of-the-art m2m alignment tool. In all our experiments we have taken a strictly language independent approach. Each of the language pairs were processed automatically with no special treatment. 1

Read the paper · More papers on PaperTik