Design and Develop A Part of Speech Tagging for Ge’ez Language using Deep Learning Approach

Asnak Yihunie Kassahun, Tessfu Geteye Fantaye · 2022

Part of Speech (POS) tagging is one of the basic and important applications of Natural Language Processing (NLP).Different POS tagging systems in various languages have been developed. In some languages, POS tagging works well with higher accuracy, but in the Ge’ez language, it is still an unsolved problem. Ge’ez is a morphological rich, free word order language, and very ambiguous where every word has many more variants based on its suffixes and prefixes. This paper proposed a deep learning-based POS tagging for Ge’ez language. The Gated Recurrent Unit(GRU), Bidirectional Gated Recurrent Unit(Bi-GRU), Long-Short Term Memory(LSTM), and Bidirectional Long-Short Term Memory(Bi-LSTM) models were developed by varying the number of training epochs and important hyperparameters. Ge’ez language has not standard POS corpus. Therefore, we have taken 2,552 sentences from Holy Bible then divided into 1,786 as a training data set and 766 as a test data set. The corpus has 26,607 total words and the experimental results show that the Bi-GRU model has achieved an accuracy value of 86.70% with 10 epochs,2 hidden layers, 128 neurons per hidden layer, and a learning rate of 0.01. The second best model is the GRU model with an accuracy value of 86.15% using the same hyperparameters with Bi-GRU model. The final result shown in the experiment shows that the GRU and Bi-GRU Ge’ez POS tagging has performed better according to tag-wise classification results shown in and the overall accuracy results of the training and testing data set. This implies that the GRU and Bi-GRU based POS tagger outperforms the LSTM and Bi-LSTM tagger when used separately. Overall, deep learning approach is better than other traditional approaches for developing POS tagger for low-resource language-Ge’ez.

Read the paper · More papers on PaperTik