Parts of Speech Tagging for Handwritten Sindhi Sentence using Deep Learning Models

Marya Soomro, Muhammad Ahsan Raza Mughal, Muhammad Khalid Sheikh, Azhar Ali Shah · University of Sindh Journal of Information and Communication Technology · 2026

Part-of-Speech (POS) tagging is an essential task in Natural Language Processing (NLP) that helps identify the role of each word in a sentence. For handwritten low-resource languages, this task becomes more difficult because of the lack of available datasets and differences in individual writing styles. In this study, a deep learning-based approach is developed for POS tagging of handwritten Sindhi sentences. For this purpose, a dataset of handwritten Sindhi sentences was collected and manually labeled with corresponding POS categories. The prepared dataset was then used for training and evaluating the proposed models. The study applies and compares two deep learning models, Long Short-Term Memory (LSTM) and Bidirectional Encoder Representations from Transformers (BERT), for automatic POS tagging. The models were evaluated using accuracy, precision, recall, and F1-score metrics. Experimental results demonstrate that BERT achieved better performance compared to LSTM due to its ability to capture contextual information from sentence structures. BERT achieved an accuracy of 87.3% while LSTM achieved 85.1%. The proposed work contributes to Sindhi language processing by providing a handwritten dataset and a baseline deep learning approach for POS tagging of a low-resource language.

Read the paper · More papers on PaperTik