Correcting Auditory Spelling Mistakes in Jordanian Dialect Using Machine Learning Techniques
Malak Smadi, Gheith A. Abandah · 2024
This paper explores the application of machine learning techniques, specifically leveraging an efficient transformer model known as Byt5, to rectify Arabic auditory spelling mistakes. The model is trained using datasets with synthetic errors generated through stochastic error injection. Evaluation is conducted using a test dataset that has auditory spelling mistakes which is gathered from Jordanian accounts on social media platforms. Results indicate that the model achieves high error correction accuracy rates, with the best performance achieved when the model is trained on a dataset with specific error injection rate (EIR). This model achieves the best character error rate of 1.13% on the test set when trained with EIR of 75%. These findings underscore the effectiveness of employing machine learning, particularly pre-trained transformer models, in addressing Arabic spelling errors, showcasing potential applications in Arabic language processing tasks.