Out-of-Vocabulary Handling in Part-of-Speech Tagging

Muhammad Alfian, Umi Laili Yuhana, Daniel Siahaan, Harum Munazharoh, Eric Pardede · International Journal on Semantic Web and Information Systems · 2025

Part-of-speech (POS) tagging is a key preprocessing step for many NLP tasks. Its broad use in education links this study to Sustainable Development Goals in quality education. Yet models still struggle with out-of-vocabulary (OOV) words. This review maps current solutions. The authors screened 1,357 papers (Jan 2014–Jun 2024) from six databases—Mendeley, IEEE Xplore, ACM DL, SpringerLink, ScienceDirect, and ProQuest—and retained 50 high-quality studies. The review shows that the field has entered a maturity phase, with established approaches—including preprocessing strategies, hand-crafted features, and learned features—being used to address OOV words in POS tagging. Nevertheless, challenges remain in coping with language variation and low-resource data, which indirectly affect model accuracy. This review provides insights into how semantic-web-based strategies can be integrated to overcome these issues and offers guidance not only for POS tagging but also for other NLP tasks involving rare or unseen words.

Read the paper · More papers on PaperTik