A Hybrid Approach for Auto-Correcting Grammatical Errors Generated by Non-Native Arabic Speakers
Zainab Althafir, Rawan Ghnemat · 2022
Spelling correction is among the most substantial Natural Language Processing (NLP) tasks, which is used as a pre-or post-processing step in many other tasks such as Optical Character Recognition (OCR). Many challenges may face while attempting to implement correctors, including correcting the real-word errors in which the wrong word is one of the language vocabularies; however, its appearance in the context is senseless. The non-native speakers who learn Arabic as a second language might have various spelling errors, especially grammatical errors involving feminine-masculine and definite-indefinite, which are not common among native Arabic speakers; thus, there is a need to employ a spell corrector that can handle such errors however the available spell correctors are not efficient enough to work with their mistakes. The proposed approach employed a rule-based system and the Arabic Bidirectional Encoder Representations from Transformers (AraBERT) model to implement an Arabic grammatical auto-corrector for non-native speakers using the Qatar Arabic Language Bank (QALB) corpus. The corpus comprises 622 sentences containing grammatical and spelling errors generated by non-native speakers. The proposed approach has enhanced the previous works by more than 17%, in which the f1-score was 45.68%. The rules, which are the main contribution, had handled about 30%, and the AraBERT model corrected the rest.