Poster: A Novel Approach for POS Tagging of Pashto Language
Haris Ali Khan, Muhammad Junaid Ali, Umm E Hanni · 2020
Pashto is a language that belongs to the Indo-European family, mostly spoken in South Asian countries, especially in Pakistan and Afghanistan. To build software that enables us to translate Pashto sentences into various languages and building Natural Language Understand (NLU) applications to make interactive Pashto software requires a well-defined corpus and Parts of Speech (POS) tagging approach. Therefore, a well-defined corpus is developed by scraping data from different websites. We have prepared dataset according to the guidelines written for Persian and Arabic languages, as these languages are somehow similar to these languages. Training of POS tagging using BiLSTM with GloVe embedding shows the effectiveness of our proposed approach and achieved 97% accuracy.