A Large-Scale Bengali Paraphrase Generation Dataset: A Comprehensive Study of Language Models and Performance Analysis

Md Mehrab Hossain, Nanziba Khan Biva, Nafiz Nahid, Ashraful Islam · 2024

This work explores the field of natural language processing with an eye toward Bengali difficulties with paraphrasing. The output offers a sizable synthetic dataset for paraphrase generation with 700,000 pairs. Along with different review sentences, the dataset includes formal sentences such as those found in newspapers as well as informal sentences from social media sites, therefore setting a new benchmark and the dataset distribution consists of 573,800 and 126,200 pairs correspondingly. Moreover, by combining a multilingual pivoting strategy, and a thorough filtering process applied to a Bengali paraphrase generation corpus, this approach promises unmatched diversity and semantic consistency compared to current available datasets. Refined many paraphrasing producing language models and analyzed them against current models and datasets in order to improve model performance study. Thus, models trained on the new dataset outperformed performance in every test category in most assessment metrics, including BLUE, ScareBLEu, ROUGE, PINC, and BERTScore. Through the provision of an enhanced dataset, novel models, and perceptive evaluation of the potential and challenges of natural language comprehension in Bengali, this work represents a noteworthy advancement in Bengali language processing, hence augmenting the effectiveness of NLP models for this language.

Read the paper · More papers on PaperTik