Findings of the WMT 2020 Shared Task on Parallel Corpus Filtering and Alignment
Philipp Koehn, Vishrav Chaudhary, Ahmed El-Kishky, Naman Goyal, Peng‐Jen Chen, Francisco Guzmán · 2020
Following two preceding WMT Shared Tasks on Parallel Corpus Filtering (Koehn et al., 2018(Koehn et al., , 2019)), we posed again the challenge of assigning sentence-level quality scores for very noisy corpora of sentence pairs crawled from the web, with the goal of sub-selecting the highest-quality data to be used to train machine translation systems.This year, the task tackled the low resource condition of Pashto-English and Khmer-English and also included the challenge of sentence alignment from document pairs.10 participants from companies, national research labs, and universities participated in this task.