Enhancing Bengali UPOS Tagging via Semi-Supervised Learning with BERT and Confidence-Based Pseudo-Labeling

Radha Rani Paul, Md. Ali Hossain · 2025

The Universal Part-of-Speech (UPOS) tag classification problem is crucial for linguistic applications and is one of the most significant tasks in Natural Language Processing (NLP). This study utilizes a semi-supervised learning strategy associated with a classification method for Bengali UPOS tagging to ensure abundance by leveraging a fine-tuned BERT model with$k$-fold cross-validation, utilizing both labeled and unlabeled data. In order to guarantee robustness utilizing supervised learning techniques, we first used a multilingual BERT model that is fine-tuned using k-fold cross-validation on a gold-labeled Bengali dataset. Subsequently, a pseudolabeling strategy is applied to unlabeled Bengali data, and high-confidence predicted UPOS tags are selectively incorporated. The resulting augmented dataset, comprising both manually labeled and filtered pseudolabeled samples, is used to further retrain the model, thus significantly boosting tagging performance and achieving an F1 score from 0.96 to 0.9951.

Read the paper · More papers on PaperTik