A discriminative approach to filter out noisy sentence pairs from bilingual corpora

Kaveh Taghipour, Nasim Afhami, Shahram Khadivi, Saeed Shiry Ghidary · 2010

Parallel corpora are essential for training statistical machine translation models. Since parallel sentence-aligned corpora are usually noisy due to inexact automatic methods when generated from parallel or comparable documents, we need to clean parallel corpora. In this paper, new features are introduced to assess the correctness of a sentence pair. Also, the impact of new features in combination with state-of-the-art features introduced in the literature is systematically evaluated. Statistical methods have been used for feature extraction and therefore this approach is independent to language. In order to better understand the problem characteristics, four supervised classification algorithms are used to classify sentence pairs as noise or parallel. Evaluating the models by taking accuracy and f-measure into account shows that using the system for cleaning a noisy parallel Farsi-English corpus, the maximum entropy model performs better than the main filtering techniques used in this paper and shows a significant improvement over two other systems.

Read the paper · More papers on PaperTik