QCMUQ$@$QALB-2015 Shared Task: Combining Character level MT and Error-tolerant Finite-State Recognition for Arabic Spelling Correction
Houda Bouamor, Hassan Sajjad, Nadir Durrani, Kemal Oflazer · 2015
We describe the CMU-Q and QCRI's joint efforts in building a spelling correction system for Arabic in the QALB 2015 Shared Task.Our system is based on a hybrid pipeline that combines rule-based linguistic techniques with statistical methods using language modeling and machine translation, as well as an error-tolerant finite-state automata method.We trained and tested our spelling corrector using the dataset provided by the shared task organizers.Our system outperforms the baseline system and yeilds better correction quality with an F-score of 68.12 on L1test-2015 testset and 38.90 on the L2-test-2015.This ranks us 2nd in the L2 subtask and 5th in the L1 subtask.