The Significance of Global Vectors Representation in Sarcasm Analysis

Christopher Ifeanyi Eke, Azah Anir Norman, Liyana Mohd Shuib, Faith B. Fatokun, Isaiah Michael Omame · 2020

Sarcasm detection is an unending drawback in natural language processing, which brings a hindrance in a progression of finding the correct state of people's sentiment. Many approaches has been applied for sarcasm detection in text such as machine learning based, lexicon based, and deep learning based approach by employing different word embedding representation schemes. The traditional word embedding schemes for learning vector representation such as bag-of-words, skip gram and Continuous bag-of-word has successfully done well in extracting discriminative syntactic and semantic features by employing vector arithmetic. However, these models failed to capture the word context in the text. In addition, skip gram and CBOW do not work openly on corpus co-occurrence statistics but rather examine the context window over the whole corpus, which fail to consider the important recurrence in the data. In order to address the aforementioned limitations, this study investigates the word vector model (GloVe) that integrate two major model families; Global statistics of matrix factorization technique and local context window based approach. The model is able to build a word representation that do not only learn semantics and grammar information, but also captures the context of the word and information of global corpus statistics. We analyzed the predictive performance of the representation by investigating sarcasm identification in tweets by employing a machine learning approach. The experiment results produced a better accuracy, which shows the significance of GloVe embedding in sarcasm classification.

Read the paper · More papers on PaperTik