UMMU$@$QALB-2015 Shared Task: Character and Word level SMT pipeline for Automatic Error Correction of Arabic Text
Fethi Bougares, Houda Bouamor · 2015
In this paper we present the LIUM (Laboratoire d'Informatique de l'Universit du Maine) and CMU-Q (Carnegie Mellon University in Qatar) joint submission in the Arabic shared task on automatic spelling error correction.Our best system is a sequential combination of two statistical machine translation systems (SMT) trained on top of the MADAMIRA output.The first is a Character-based one, used to produce a first correction at the character level.Characters are then glued to form the input to the second system working at the Word level.This sequential combination achieves an F 1 score of (69.42) that is better than the best F 1 score reported on the 2014 test set (67.91).The UMMU best submission to the QALB-15 shared task is ranked first over 10 submission on the L2 test condition and second over 12 submission on the L1 testsset.