Pashto Shallow Parsing: A Deep Learning Approach
Ajmal Samadi, D. Muhammad Noorul Mubarak, Noor Ahmad Emall · 2025
This paper presents the first deep learning-based shallow parsing system for the Pashto language, addressing the significant lack of syntactic tools for this low-resource and morphologically rich language. A comprehensive corpus of over 15,000 manually annotated sentences was developed using texts from reliable Pashto news sources. Each token was labeled using the IOB format, covering major phrase types such as noun, verb, prepositional, and adverbial phrases. Leveraging this resource, a Bidirectional Long Short-Term Memory (BLSTM) network was designed and trained to capture the syntactic structure of Pashto sentences. The proposed model achieved a test accuracy of 98%, demonstrating strong generalization and effective learning of phrase boundaries. The experimental results confirm the ability of the model to handle complex sentence structures and the key characteristics of Pashto. This work not only contributes a high-quality annotated dataset but also establishes a robust baseline for future syntactic and semantic processing tasks in underrepresented languages. Both the dataset and trained model will be publicly released to facilitate further advancements in Pashto NLP.