Neural Methods for Aligning Large-Scale Parallel Corpora from the Web for South and East Asian Languages
Philipp Koehn · 2024
We introduce neural methods and a toxicity filtering step to the hierarchical web mining approach of Paracrawl (Bañón et al., 2020), showing large improvements.We apply these methods to web-scale parallel corpus mining for 9 South and East Asian national languages, creating training resources for machine translation that yield better translation quality for most of these languages than existing publicly available datasets in OPUS.Our methods also generally lead to better results than the global mining approach of Schwenk et al. (2021).