Pre-Trained Word Embeddings for Sarcasm Detection in Indonesian Tweets: A Comparative Study

Mochamad Alfan Rosid, Daniel Fernando Siahaan, Ahmad Saikhu · 2022

In affective computing, sarcasm detection is vital because sarcasm can affect the polarity of sentiment analysis. Sarcasm is one of the most challenging problems researchers face when conducting sentiment analysis. Sarcasm is difficult to identify in text because the meaning of the words expressed by a person are the opposite to what the person really means. Deep learning is currently widely used for sarcasm detection. For generating vector representations of words, three distinct pre-trained word embedding models were employed in this study, namely GloVe, fastText, and BERT. Furthermore, three distinct deep learning architectures were utilized, namely Bidirectional Long Sort-Term Memory (BiLSTM), Bidirectional Gated Recurrent Unit (BiGRU), and Convolutional Neural Network (CNN), for sentence-level sarcasm detection in tweets written in Indonesian. The dataset was collected through data crawling using Twitter API. The data was then further preprocessed and features were extracted. The experimental results indicate that the combination of fastText embeddings and BiGRU as the classifier produced the best performance, with an accuracy of 93.85%.

Read the paper · More papers on PaperTik