Binary Encoding Based Morpheme Boundary Detection of Dogri Language
Parul Gupta, Shubhnandan Singh Jamwal · International Journal of Intelligent Engineering Informatics · 2024
Machine learning (ML) models like decision tree, SVM, random forest and KNN are generally used with structured data of morpheme boundary detection but binary representation directly presents the data in a format that can be used by the models without pre-processing. Dogri is an Indo-Aryan language spoken primarily in the state of Jammu and Kashmir, as well as in certain regions of the neighbouring states of Himachal Pradesh and Punjab. This research paper explores and analyses common ML models which are rarely applied in detecting morpheme boundaries. The dataset of 10,000 Dogri words along with their morpheme boundaries in bit values are used for training and evaluation. In this paper, we trained the bi-LSTM and ML models on a different dataset and observed that bi-LSTM outperformed other ML models and exhibited a remarkable recall of 69.50%, 79.64%, and 81.33% on three different datasets respectively.