Partitioning parallel documents using binary segmentation

Jia Xu, Richard Zens, Hermann Ney · 2006

In statistical machine translation, large numbers of parallel sentences are required to train the model parameters. However, plenty of the bilingual language resources available on web are aligned only at the document level. To exploit this data, we have to extract the bilingual sentences from these documents.

Read the paper · More papers on PaperTik