Comparison of Turkish Paraphrase Generation Models
Ahmet Bagci, Mehmet Fatih Amasyalı · 2021 International Conference on INnovations in Intelligent SysTems and Applications (INISTA) · 2021
Paraphrase generation is an important NLP task of generation a sentence that has the same meaning as a given sentence. But, there are few studies on Turkish in this field. In this study, we present the most comprehensive study on Turkish. We focused on question sentences due to the resources we could access. First of all, since there is not enough large parallel paraphrase dataset for Turkish, we created several datasets with various methods. We compared the success of these datasets using the BERT2BERT architecture. We found the combination of automatically generated and manually generated datasets to be the most successful dataset. Using this dataset, we compared the success of the most popular architectures in the field. We found that MBart is the most successful model according to BLEU/Rouge criteria and BERT2BERT according to human evaluation.