Statistical and Deep Learning Approach for Kannada Parts of Speech Tagger
Akarsh Ramachandra Hegde, Sharma Akash, Adarsh Ishwar Hegde, Adarsh Nayak S R, H. R. Mamatha · 2023
With over 64 million native speakers, Kannada is one of the major languages of the Dravidian Language family in India. However, it confronts considerable difficulties when it comes to linguistic resources for processing and computing. Basic NLP operations like lemmatization, parts of speech tagging, machine-translation, sentiment analysis and recognizing named entities tasks are greatly hindered by its complex morphology and agglutinative nature. In our research work, we aim to obtain the best results in parts of speech tagging for Kannada. We achieve this by working on both CRF and Bi-LSTM, along with two widely used word embedding techniques: Word2Vec(implemented using Gensim) and IndicNLP. By meticulously applying these methodologies to two distinct datasets, we have produced a total of six distinctive models. Our research deeply explores the contrast and comparison of these models, leading us to identify the most proficient model among them. This exhaustive analysis culminates in the discovery that the Bi-LSTM model attains an impressive F1-score of 0.921 and test accuracy of 86.2%. These outcomes are primarily drawn from the primary dataset sourced from Siva Reddy, a renowned author in this field.