Experimentation Of Various Preprocessing Pipelines For Sentiment Analysis On Twitter Data About New Indonesia’s Capital City Using SVM And CNN

Siska Pebiana, Nuraisa Novia Hidayati, Dian Isnaeni Nurul Afra, Elvira Nurfadhilah, Harnum Annisa Prafitia, Junanto Prihantoro, Radhiyatul Fajri, Mohammad Teduh Uliniansyah, Agung Santosa, Lyla Ruslana Aini, Yosi Sahreza, Aulia Haritsuddin Karisma Muhammad Subekti, Josua Geovani Pinem, Muhammad Reza Alfin, Agung Septadi, Siti Shaleha, Gembong Satrio Wibowanto, Asril Jarin, Gunarso, Andi Djalal Latief · 2022

Selecting the sequence of pre-processing stages to obtain the best data for subsequent processes is still challenging. We have conducted a series of experiments with various sequences of pre-processing methods to acquire the best data. For the experiments, we used 12. 5K Indonesian Twitter on the new capital city of Indonesia (IKN) and 16. 2M monolingual text. We compared the manual labeling results with the outputs of SVM and CNN, which use the word-embedding feature generated by Word2Vec. The proposed pre-processing pipeline can be used as a reference for sentiment analysis research using Indonesian social media data. The best experimental results show a better F-1 value in SVM with B.2 pipeline (normalized word, stop word) is 69.32% and CNN with B.3 pipeline (normalized word, stop word, stemming) is 73.03%.

Read the paper · More papers on PaperTik