OpenITI-proc corpus
Yonatan Belinkov, Alexander Magidow, Alberto Barrón‐Cedeño, Avi Shmidman, Maxim Romanov · Zenodo (CERN European Organization for Nuclear Research) · 2019
Arabic is a widely-spoken language with a long and rich history, but existing corpora and language technology focus mostly on modern Arabic and its varieties. This is a large-scale historical corpus of the written Arabic language, spanning 1400 years. The corpus is processed with the Farasa Arabic NLP toolkit. The corpus can be used to study the history of the Arabic language.