Improving the Accuracy of Semantic Similarity Prediction of Arabic Questions Using Data Augmentation and Ensemble
Shorouq M. AlAwawdeh, Gheith A. Abandah · 2021
Finding the semantic similarity of a pair of questions is an important application of Arabic natural language processing, e.g., for providing an answer to a new question similar to another question with a known answer. Modern pre-trained models are currently re-tasked for this application. We present here proposed solutions to improve the accuracy of such models. We investigate training the acclaimed pre-trained AraBERT model in two phases, using data augmentation, and using ensemble of predictors. Two-phase training and data augmentation are used to provide more data examples for the model to learn and generalize. The proposed solution improves the error rate from the best previous work of 4.08% to 3.12% on the benchmark Mawdoo3 dataset. The proposed model is first trained on a back-translation version of the dataset, then it is retrained in a second phase on the dataset with transitive, symmetric, and reflexive data augmentation methods. The prediction accuracy is further improved by constructing an ensemble of five such trained models and selecting the majority vote.