Towards Developing an Automatic Punctuation Prediction Model for Bangla Language: A Pre-trained Mono-lingual Transformer-based Approach
Ayman Iktidar, Hasan Murad · 2024
Punctuation prediction is an essential task in natural language processing that involves determining the appropriate position of punctuation marks in a given text. Accurate punctuation contributes to better readability and understanding in the written form of a language. A significant amount of previous research work on punctuation prediction has been found in resource-enriched languages like English. For automatic punctuation prediction in Bangla, multilingual pre-trained transformer models have been used. However, the Bangla language has unique grammatical features that make automatic punctuation prediction challenging using multi-lingual pre-trained transformer models. In this research work, we have proposed an automatic punctuation prediction approach for the Bangla language using a mono-lingual pre-trained transformer model specifically designed for Bangla. By modifying an existing dataset with four punctuation classes, a new Bangla dataset has been developed with five punctuation classes: Comma, Period, Question, Exclamation, and Others. Using the new dataset, we have trained and evaluated different versions of mono-lingual transformer-based models for automatic punctuation prediction in Bangla. We have found the state-of-the-art result for Bangla using BanglaBert-large with an F1 score of 0.89 for the existing dataset. Our model has attained an F1 score of 0.74 across the five punctuation classes providing state-of-the-art results.