Arabic Part Of Speech (POS) Tagging Analysis using HMM Trigram method on Al-Qur’an Ayah Sentences

Arief Fatchul Huda, Izki Zakiyah Al-Hamro, Asep Solih Awalluddin, Muhammad Ibnu Pamungkas · 2021

Part Of Speech (POS) tagging is part of Natural Language Processing to determine correctly the label in a sentence from the given input. Different POS tagging techniques in some literatures have been developed for English text, and few for Arabic texts. This problem uses a method based on the second hidden Markov model, which is looking for to two words to the past or better known as the HMM Trigram method. The main problems of POS tagging are Out Of Vocabulary (OOV) and word ambiguity. This study discusses POS tagging using HMM Trigram method on Al-Qur'an text data. The dataset is divided into three categories of data derived from quran corpus consists of 150 simple perfect sentences, 50 sentences with more than one S/P/O/K and 50 selected verses of the Qur'an. The data experiment was carried out using a cross-validation technique, namely k-fold cross validation. The data is classified into two, namely training data and test data. Training data is used to find emissions and transitions probabilities, while data testing uses the Viterbi algorithm. The experimental results achieved an average accuracy of 86% for simple datasets, 60% for medium datasets, and 38% for the complete paragraph datasets.

Read the paper · More papers on PaperTik