A Comparative Study on Word Embeddings and Social NLP Tasks

Fatma Elsafoury, Steven R. Wilson, Naeem Ramzan · 2022

In recent years, grey social media platforms, those with a loose moderation policy on cyberbullying, have been attracting more users.Recently, data collected from these types of platforms have been used to pre-train word embeddings (social-media-based), yet these word embeddings have not been investigated for social NLP related tasks.In this paper, we carried out a comparative study between social-mediabased and non-social-media-based word embeddings on two social NLP tasks: Detecting cyberbullying and Measuring social bias.Our results show that using social-media-based word embeddings as input features, rather than non-social-media-based embeddings, leads to better cyberbullying detection performance.We also show that some word embeddings are more useful than others for categorizing offensive words.However, we do not find strong evidence that certain word embeddings will necessarily work best when identifying certain categories of cyberbullying within our datasets.Finally, We show even though most of the state-of-the-art bias metrics ranked social-media-based word embeddings as the most socially biased, these results remain inconclusive and further research is required.Content Warning: As part of our experiments, we show some offensive words.* Throughout this paper, we differentiate between the terms "offensive" and "profane": we use the term "offensive" to describe an expression that is offensive to a group of people but not necessarily profane e.g."women belong to the kitchen" while we use the term "profane" to describe expressions like "b*tch".

Read the paper · More papers on PaperTik