Partitioning parallel documents using binary segmentation
Jia Xu, Richard Zens, Hermann Ney · 2006
In statistical machine translation, large numbers of parallel sentences are required to train the model parameters. However, plenty of the bilingual language resources available on web are aligned only at the document level. To exploit this data, we have to extract the bilingual sentences from these documents.