A Gold Multipurpose Arabic Corpus (GAC)

Hussein Awdeh, Adelle Abdallah, Youssef Zaki, Gilles Bernard, Mohammad Hajjar · 2021 International Conference on Electrical, Computer and Energy Technologies (ICECET) · 2021

A corpus is a large collection of spoken or written one or more language, it is collected from different resources, then structured, stored, and treated automatically by special algorithm to reach the desired goal, this is why it is used by researchers in different domain such as grammar, semantics, lexicography, natural Language Processing and other language studies. Therefore, building a corpus was still a challenge for many researchers in those fields for many years. Our study in this paper aims to build a new gold standard Arabic corpus due to the lack of successful trials in compiling Arabic corpora. The corpus produced by our team, is a text corpus, collected from a set of Arabic Newspaper articles morphologically analyzed from eight Arabic countries, and it contains more than 18 million words in total, covering six categories (Religion, Economy, Culture, Sports, Local and International News). It was encoded with UTF -8 encoding and marked with two mark-up languages: JSON and XML. In the hope that this corpus can be used as an accurate reference for segmentation and validation and learning in the syntax analysis mainly for the word segmentation and part of speech tagging.

Read the paper · More papers on PaperTik