Design and Develop Automatic Parts of Speech Tagger for Siltigna Language
Biniyam Sime Dana, Michael Melese Woldeyohannis, Seiyfedin Mohammed Yesuf, Eyobed Birhanu Paulos, Mesay Gemeda Yigezu · 2024
The purpose of this study is to design and develop automatic part-of-speech taggers for the Siltigna language using the deep learning algorithm. Part of speech tagging is the process of assigning words within sentences to the appropriate word categories. In most languages POS tagger is utilized as a pre-processing component in several NLP applications. Considering that several part-of-speech taggers are developed for various foreign and local languages. However, those POST cannot be applied directly with Siltigna due to morphological difference between the languages. To build tagging modelsvarious deep learning algorithms such as Standard RNN, GRU, LSTM, and BiLSTM were implemented. For training and testing purposes, 3500 sentences were collected from different sources, such as Siltigna text books, annual and monthly reports from the tourism o ce, and Siltigna newspapers, in sound format, then converted to text format. Some proverbs, which are written in Siltigna, were also collected to balance the sample corpus, and 22 sets of tags were identified for tagging purposes. The POST tagger is tested using a percentage spilled mechanism. The experimental result indicates that for simple RNN, LSTM, GRU and bidirectional LSTM was 92. 49 %, 92. 55 %, 92. 61 % and 93. 05 % test accuracy, respectively. Based on the result obtained from the accuracy test, conclusions and recommendations are forwarded.