Development of Voting-Based POS Tagger for URDU Language

Usama Ahmed, Ahmed Raza, Kainat Saleem, Amna Sarwar · Journal of Computational Science and Applications (JCSA) ISSN 3079-0867 (Onilne) · 2025

The process of sequence labeling (POS) by assigning syntactic tags to words in the given context plays an important role in various NLP applications. The core motive of this work is to tackle the morpho-syntactic category of words in Urdu. This language has many computational challenges because of its dual nature, as different tags propose different morpho-syntactic rules. Our work comprises different tasks; initially, we tracked the best combination of feature sets in terms of CRF to enhance previous results on two stable and well-known datasets, the Bushra Jawaid dataset and the CLE dataset. Due to syntactic ambiguity, a state-of-the-art voting method has been introduced to overcome contradictory results from different machine learning classifiers. The results show significant improvement over baseline results, with an F1-score of 94.8% on the primary dataset and 95.7% on the succeeding dataset. Our work also incorporates a deep learning model, Long Short-Term Memory (LSTM), for one of the most diverse and inflectional tasks, Part of Speech Tagging for the Urdu language, achieving F1-scores of 86.7% and 96.1% on both datasets, respectively.

Read the paper · More papers on PaperTik