The International Corpus of Arabic: Compilation, Analysis and Evaluation

Sameh Alansary, Magdy H. Nagi · 2014

This paper focuses on a project for building the first International Corpus of Arabic (ICA). It is planned to contain 100 million analyzed tokens with an interface which al-lows users to interact with the corpus data in a number of ways [ICA website]. ICA is a representative corpus of Arabic that has been initiated in 2006, it is intended to cover the Modern Standard Arabic (MSA) language as being used all over the Arab world. ICA has been analyzed by Bibliotheca Alexandrina Morphological Analysis Enhancer (BAM-AE). BAMAE is based on Buckwalter Arabic

Read the paper · More papers on PaperTik