Parts of Speech Tagging in Bangla Sentences using Supervised Learning: A Performance Comparison between Viterbi and Bidirectional-LSTM Models
Mosarrat Rumman, Abu Nayeem Tasneem, Md. Golam Rabiul Alam · 2021
Parts of speech (POS) tagging is a crucial preprocessing step for many Natural Language Processing applications. Though numerous works have been done on English corpus with high accuracy, very few works have been done on Bangla Corpus due to scarcity of resources and the ambiguity of the language. In this paper we have created a POS tagger using Hidden Markov Model(HMM) with Viterbi Algorithm for decoding and a deep learning model called Bidirectional Long-short term memory (BiLSTM). We used similar datasets to compare the performance of the two approaches. It can be inferred from the results that increasing the size of dataset has greater positive impact on the performace of Bi-LSTM model than on the HMM model.