Building Arabic corpora from Wikisource
Imene Bensalem, Salim Chıkhı, Paolo Rosso · 2013
This paper describes a new tool that helps extracting clean text from the Arabic Wikisource dump in order to build corpora. The tool purpose is illustrated by the generation of a subcorpus from Wikisource, which is a step towards the building of an evaluation corpus for Arabic intrinsic plagiarism detection.