Building a Part of Speech tagger for the Tamil Language
Kengatharaiyer Sarveswaran, Gihan Dias · 2021
Identifying the lexical category or Part of Speech (POS) of words is critical for developing Natural Language Processing (NLP) systems. Existing Tamil POS taggers are either not publicly available, inaccurate or use a non-standard POS tagset. In this paper, we present a state-of-the-art, contextual neural POS tagger named ThamizhiPOSt. It tags words in a sentence with their lexical category, using the Universal Part of Speech tagset. Each tag is based on the word's context in the sentence. ThamizhiPOSt also integrates a tokeniser and a sentence segmenter so that it can handle raw Tamil text. This paper introduces this POS tagger and compares it with existing tools and resources for Tamil POS tagging. ThamizhiPOSt reports an accuracy of 93.27% for unseen data, which constitutes the best score in comparison to the publicly available Tamil POS taggers. ThamizhiPOSt and its associated resources are all made public. Further, ThamizhiPOSt has been published as a Python library for others to build upon.