Improving Hidden Markov Model for very low resource languages: An analysis for Assamese parts of speech tagging
Diganta Baishya, Rupam Baruah · 2021
Recent advances in research in the field of natural language processing have facilitated many new applications. Natural Language Processing basically involves automating the synthesis and generation of natural languages. However, most of the developments have happened only for a few dominant languages spoken widely like English, Chinese, etc. Some of the developments are also observed for dominant Indian languages like "Hindi", "Bengali" Tamil, etc. However, research into most of the other languages spoken in the world is at a very primitive stage. This paper aims to highlight issues related to one such language for which many resources are not available in a computable format. We report our initial work on Assamese with limited resources. Assamese is the official language of Assam, India and it is also the most spoken Indian language in North East India. We report results of a few experiments carried with a very low amount of training, as resources available for Assamese are not adequate. Part of speech tagging is fundamental to any NLP application. Though some amount of research has been conducted for automatic parts of speech tagging, they have used large amount of data that are not available in public domain. We have modified the viterbi algorithm used for Hiden Markov Model and applied it for automatic parts of speech tagging for Assamese. We have also applied certain language characteristics that can be used for improving accuracy. The key is to use a training dataset as small as possible so that the dependency on training data is limited. We conclude the paper with a brief discussion on the scope of Assamese POS tagging in the future.