Toward Better Chinese Word Segmentation for SMT via Bilingual Constraints

Xiaodong Zeng, Lidia Sam Chao, Derek F. Wong, Isabel M. Trancoso, Liang Tian · 2014

This study investigates on building a better Chinese word segmentation model for statistical machine translation.It aims at leveraging word boundary information, automatically learned by bilingual character-based alignments, to induce a preferable segmentation model.We propose dealing with the induced word boundaries as soft constraints to bias the continuous learning of a supervised CRFs model, trained by the treebank data (labeled), on the bilingual data (unlabeled).The induced word boundary information is encoded as a graph propagation constraint.The constrained model induction is accomplished by using posterior regularization algorithm.The experiments on a Chinese-to-English machine translation task reveal that the proposed model can bring positive segmentation effects to translation quality.

Read the paper · More papers on PaperTik