Arabic learner corpus (ALC) v2: a new written and spoken corpus of Arabic learners
Abdullah Ahmad M. Alfaifi, Eric Atwell, Hedaya Ibraheem · White Rose Research Online (University of Leeds, The University of Sheffield, University of York) · 2014
Arabic learner corpora have not received enough attention, particularly for learning Arabic as a second language (in Arabic speaking countries). Based on the literature, there are a few projects are developing Arabic Iearner corpora, of which most are not freely available for users or researchers. In addition to that they are intended to assist in the language acquisition of Arabic as a foreign language(collected from learners studying Arabic in non・Arabic speaking countries). The present paper aims to introduce the Arabic Learner Corpus. It is being developed at Leeds University, and oomprises of 282,732 words, collected from learners of Arabic in SaudiArabia. The corpus includes written and spoken data produced by 942 students, from 67 different nationalities studying at pre-university and university levels. The paper focuses on two angles of this corpus, the design criteria and the content. The design criteria of the ALC discuss the target language, the participants, the corpus size, the materials included, the method of data collection, the metsdata of the corpus materials and contributors, and text distribution. The second part, ALC contents, is illustrated based on 26 elements representing the corpus metsdata. The goal of the ALC is to provide an open-source of data for some linguistic researeh areas related to Arabic language learning and teaching. So, the corpus data is available for download in TXT and XML formats, hand-written sheets which are in PDF format as well as the audio recordings which are available in MP3 format.