An Empirical Analysis of PoS Tagging for Kannada Machine Translation

Jamuna Jamuna, H. R. Mamatha · 2023

Parts of Speech (POS) tagging process has emerged as one of the crucial and very basic preprocessing technique for any natural language processing tasks. In Kannada, a dominant language in southern India, morphologically rich and low resourced at the same time, PoS tagging process was difficult to achieve in the beginning. Later, studies show that Conditional Random Fields, Hidden Morkov Model and deep learning techniques have produced good accuracy. This paper investigates all the three above mentioned models and argues that deep learning model, which uses bidirectional Long Short Term Memory as a RNN unit, produces the highest accuracy of 93% in contrast to CRF and HMM model with a precision accuracy of 65% and 42% respectively. Also, the paper specifies how important a PoS tagging process is in the task of Machine Translation, which is booming in the world of computational linguistics.

Read the paper · More papers on PaperTik