A Model of Preprocessing For Social Media Data Extraction

Dodo Zaenal Abidin, Siti Nurmaini, Reza Firsandaya Malik, Jasmir Jasmir, Errissya Rasywir, Yovi Pratama · 2019

Tropical disease grows fast and requires detection. One source of data for detections is social media Twitter. However, social media data has data with diverse data structures primarily from unstructured user syntax and grammar. Therefore the twitter message (tweet) must be purified by a preprocessing method involving part of speech (POS) rule. This paper proposes a preprocessing model for twitter data to obtain a clean dataset. There are some steps. Firstly, we use Out-of-Vocabulary (OOV) word to analyse the tweets from Indonesian texts. Secondly, in stemming step, we use Sastrawi library. We also compare the result of tokenization and combining with Out-of-vocabulary (OOV) Word, Stemming N-Gram, and Stop Word Removal Sastrawi library into a well preprocessing approach. From the experimental result, we can get the result of Preprocessing task related to tweet data characteristic in Indonesia language. We can conclude that our form get more valuable result in terms of meaningful word occurrence comparing to the result obtained by just running common preprocessing tasks.

Read the paper · More papers on PaperTik