A hybrid part-of-speech tagger for Marathi sentences
Madhuri M. Deshpande, Sharad D. Gore · 2018
With thousands of languages in the world, and the increasing speed and quantity of information being distributed across the world, automatic translation between languages by computers, Machine Translation, has become an increasingly important area of research. For a machine to translate text in one natural language to target text in another language, it requires an understanding of the language, its grammar (syntax), its meaning (semantics) and the ability to use this knowledge for making inferences. Words have definite meaning(s) making them deterministic and finite. Words are not ambiguous in their meaning. Context dependency arises when a word is used with a group of words (bag-of-words) in a specific way that causes its meaning to be dependent on the group of words. In this paper, we present a hybrid, multi-pass Part-Of-Speech (POS) tagger developed for Marathi sentences which builds feature vector for each word in a sentence by referring to the previous and next word preceding and succeeding the current word that is being tagged. The analysis of the Marathi input sentence is done first by tokenizing each word in the sentence and finding the stem for each token. Every token is analyzed for its POS tag, the tense, mood and aspect. This process is POS tagging. Ambiguities may arise in the process of tagging.