Document Alignment for Generation of English-Punjabi Comparable Corpora from Wikipedia

Vishal Goyal, Ajit Kumar, Manpreet Singh Lehal · International Journal of E-Adoption · 2020

Comparable corpora come as an alternative to parallel corpora for the languages where the parallel corpora is scarce. The efficiency of the models trained on comparable corpora is comparatively less to that of the parallel corpora however it helps to compensate much to the machine translation. In this article, the authors have explored Wikipedia as a potential source and delineated the process of alignment of documents which will be further used for the extraction of parallel data. The parallel data thus extracted will help to enhance the performance of Statistical Machine translation.

Read the paper · More papers on PaperTik