Mining parallel corpora from Wikipedia
Olivier Kraif · Studies in corpus linguistics · 2024
Abstract In this article, we address the issue of Wikipedia as a multilingual resource to extract parallel corpora that are useful in multilingual terminology extraction or machine translation. While most previous work in this field assumes that Wikipedia is suitable for mining comparable corpora, we concentrate on the actual place of translation in the editorial process of Wikipedia to examine the possibility of extracting parallel corpora, that is, texts where source segments can be linked to their translations. After identifying the different projects, tools and recommendations that allow contributors to enrich Wikipedia by exercising their skills as translators, we conduct an experiment in which we download pairs of articles containing translations. We show the importance of performing a temporal alignment of the versions to be downloaded before launching the actual sentence-level alignment. This strategy allows us to obtain a large volume of parallel texts with good-quality sentence-to-sentence alignment.