Ensemble Machine Translation to Filter Low Quality Corpus

Wuying Liu, Lin Wang · 2022

The popularization of ubiquitous computing equipment and the development of network communication technology have greatly improved the productivity of human language information. The quality of large-scale language information is a key factor affecting the effectiveness of deep learning algorithms. How to filter low quality corpus from large-scale corpus that is unstructured, non-standardized, and even contains fallacies, to obtain high quality corpus has become a research issue of great application value. We focus on the specific filtering issue of bilingual parallel corpus, re-examine the manual filtering process, and propose an ensemble machine translation filtering idea. From the perspective of translation direction, we design and implement a single-engine-based ensemble machine translation filtering framework and algorithm, and from the translation system perspective, we design and implement a multi-engine-based framework and algorithm respectively. Experimental results show that the single-engine-based algorithm can integrate the advantages of different machine translation directions, and the multi-engine-based one can integrate the advantages of different machine translation systems. Using the straightforward methods of Levenshtein string morphological similarity and linear weighted ensemble can implement efficient industrial-grade corpus filtering.

Read the paper · More papers on PaperTik