Training Doc2Vec on a Corpus of Persian Poems to Answer Thematic Similarity Multiple-Choice Questions
Soroosh Akef, Mohammad Hadi Bokaei, Hossein Sameti · 2020
This paper reports our improvement over the previous benchmark of the task of answering poetic verses' thematic similarity multiple-choice questions (MCQs). These questions, which frequently appear in the Iranian national university entrance exam, require the test-taker to select the verse which is thematically similar to the stem verse. These questions would test a natural language processing (NLP) model's ability to not only understand, but interpret a poetic verse. In this experiment, we trained a Doc2Vec model on a corpus of Persian poems and proceeded to use the trained model to get the vector representations of the poetic verses. Subsequently, the poetic verse among the options with the highest cosine similarity to the stem verse was selected as the correct answer by the model. This model managed to answer %38 of the questions correctly, which was an improvement of %6 over the previous benchmark. Furthermore, it was observed that the model's approach to answering the questions seemed different from the previous model, as there were relatively few questions that the models had both answered correctly. The ultimate goal of this task is to estimate the difficulty of such questions by using the performance of the models as a feature. The fact that the two models did not behave similarly can prove useful, as the model whose performance has the most correlation with that of a human test-taker can be used for the future task of thematic similarity MCQ difficulty estimation.