Towards Minimal Training for Low Resource Languages
Diganta Baishya, Rupam Baruah · 2023
The advancement of research in the field of natural language processing has made peoples’ daily lives much easier, with numerous applications at the disposal. The traditional methods like the Hidden Markov Model, the CRF classifier, the Naive Bayes classifier, and others are being replaced by neural networks in recent times. However, most of these methods work considerably well only with huge amounts of training data and, hence are not suitable for languages that are poor in terms of trainable resources. The challenge is to make the system work considerably well with minimal training. This paper presents research work to understand the effect of training size for part of speech tagging, which is one of the preliminary tasks for any NLP application. Experiments are conducted to understand the training size required for standard techniques to perform with high accuracy. The results of the experiments conducted for English and Assamese are presented in this paper.