Enhancing Named Entity Recognition in Low-Resource Languages: The Crucial Role of Data Sampling in Malayalam
Athira Gopalakrishnan, K. P. Soman · 2024
This paper focuses on dataset class imbalances to address Named Entity Recognition (NER) difficulties in low-resource languages like Malayalam. The main goal is to draw attention to how important data sampling is for improving NER performance. Through an exploration of various algorithms, particularly advanced models, our research demonstrates that data sampling significantly improves model precision and recall by harmonizing class distribution. Skewed datasets in low-resource languages make entity recognition less accurate, mostly because of the overrepresentation of the ‘O’ class. The work highlights how data sampling helps to progress NLP research and applications in low-resource linguistic situations and promotes successful transfer learning. The results highlight how crucial it is to correct dataset imbalances to create NER systems that are more reliable and accurate.