Hope Speech Detection on Social Media Platforms
Pranjal Aggarwal, Pasupuleti Chandana, Jagrut Nemade, Shubham Kumar Sharma, Sunil Saumya, Shankar Biradar · 2023
Since personal computers became widely available in the consumer market, the amount of harmful content on the internet has significantly expanded. In simple terms, harmful content is anything online that causes a person distress or harm. It may include hate speech, violent content, threats, Non-Hope Speech, etc. The online content must be positive, uplifting, and supportive. Over the past few years, many studies have focused on solving this problem through hate speech detection, but very few focused on identifying Hope Speech. This chapter discusses various machine learning approaches to identify a sentence as Hope Speech, Non-Hope Speech, or a Neutral sentence. The dataset used in the study contains English YouTube comments and is released as a part of the shared task “EACL-2021: Hope Speech Detection for Equality, Diversity, and Inclusion”. Initially, the dataset obtained from the shared task had three classes: Hope Speech, Non-Hope Speech, and not in English; however, upon deeper inspection, we discovered that dataset relabeling is required. A group of undergraduates was hired to help perform the entire dataset’s relabeling task. We experimented with conventional machine learning models (such as Naïve Bayes, Logistic Regression, and Support Vector Machine) and pre-trained models (such as BERT) on relabelled data. According to the experimental results, the relabelled data has achieved a better accuracy for Hope Speech identification than the original dataset. Keywords: Hope Speech, hate speech, cyberbullying, BERT, machine learning, transfer learning Contents