A Study of T’nT and CRF Based Approach for POS Tagging in Assamese Language
Ridip Ranjan Deka, Simanta Kalita, Kishore Kashyap, Manash Pratim Bhuyan, Shikhar Kr. Sarma · 2020
Assamese is one of the official languages of India which is officially recognized among 21 other official languages which is verbalized by, around 20 million people in the Northeastern part of India. In the recent time, a large number of works related to Natural Language Processing is going on for the language. Different types of models and algorithms have been employed for tagging the part of speech in a sentence, which is done by tagging each word against a predicted tag by an algorithm or by a model based on some annotated trained corpus. Assamese is a morphologically rich language. The lexical ambiguity of the words in the language is extensive. In this paper, performances of two existing tagging techniques for Assamese language have been compared, that is, Conditional Random Field and Trigrams'nTag to study the efficiency of the models in tagging of Parts-of-speeches for the language and the study might also help the Natural Language Processing (NLP) researchers in getting an understanding and analyzing the performance standard and effectiveness of this two existent models for POS tagging task for other morphologically rich Indian languages.