Informal Indonesian Part-of-Speech Tagger Using Hidden Markov Model and Normalization Algorithm

Fika Hastarita Rachman, Firdaus Solihin, Nenden Siti Fatonah, Tsania Hazhiahadani, Sri Herawati, Imamah Imamah · 2023

Twitter is a social networking platform that can be used as resources to find information in real time. Twitter is one of the social media that does not have writing criteria, users are free to express their thoughts and send messages. There are many messages and tweets that use abbreviations, foreign words, local and mixed languages, which are characteristic of informal language patterns. Informal language patterns do not rely on Indonesian, making the process of annotating word classes less precise and influential on the level of precision. The use of POS Tagging in previous research can annotate word classes well in formal sentences. The problem in informal sentences is solved by the process of word normalization after the preprocessing stage, where informal words are converted into formal words then Part-of-Speech Tagging is carried out on each word in the sentence such as nouns, verbs, adjectives, etc. The dataset used was 846 words, with a informal word count of 99 words. With the Hidden Markov Model commonly used in machine learning as well as the word normalization process, the system is able to increase the resulting precision value, there is a significant increase of up to 10% in the precision produced.

Read the paper · More papers on PaperTik