Benefits of Morphosyntactic Features on English-Arabic Statistical Machine Translation
Safae Berrichi, Azzeddine Mazroui · 2018 IEEE 5th International Congress on Information Science and Technology (CiSt) · 2018
Arabic is a Semitic language that arouses great interest in the field of machine translation. Given the large differences between morphology and grammar of Arabic and English, a statistical approach based on parallel corpus is privileged in the machine translation area. The statistical machine translation system's (SMT) quality is highly related to the preprocessing and post-processing steps of the English-Arabic parallel corpus used in the training phase. We present in this study a new alignment approach based on parallel corpus to involve some morphosyntactic information (stem, lemma and POS tags) for each word in this corpus. To evaluate our approach, we used the phrase-based statistical machine translation (PBSMT) system and then we tested it on the English-Arabic United Nations (UN) parallel corpus. The results obtained in the tests we carried out showed that this approach contributed to a significant improvement of the word alignment quality and consequently increase the BLEU score value of the translation.