Findings of the WMT 2023 Shared Task on Parallel Data Curation
Steve Sloto, Brian J. Thompson, Huda Khayrallah, Tobias Domhan, Thamme Gowda, Philipp Koehn · 2023
Building upon prior WMT shared tasks in document alignment and sentence filtering, we posed the open-ended shared task of finding the best subset of possible training data from a collection of Estonian-Lithuanian web data.Participants could focus on any portion of the end-to-end data curation pipeline, including alignment and filtering.We evaluated results based on downstream machine translation quality.We release processed Common Crawl data, along with various intermediate states from a strong baseline system, which we believe will enable future research on this topic.