A Machine Learning-Based Approach to Detect Cyberbullying Tweets

Giuseppe Prisco, Francesco Mercaldo, Antonella Santone, Mario Cesarelli · 2024

Cyberbullying has emerged as a pervasive and detrimental phenomenon and refers to the use of digital communication tools, such as social media, to harass, intimidate, or harm others. This phenomenon, coupled with the anonymity afforded by online platforms, can induces negative psychological impacts, such as heightened levels of stress, anxiety, depression, and feelings of loneliness and isolation. Therefore, the development of targeted interventions to detect offensive contents is imperative for mitigation of the harmful effects of cyberbullying and for understanding the interplay between cyberbullying and mental health. This study aimed to develop an automatic methodology capable to discriminate offensive tweets. The contents are divided in two groups: ‘non-offensive’ and ‘offensive’. Two different Feature Extraction techniques, six Machine Learning (ML) classifiers and two validation strategies were employed in this study. Great results in terms of Accuracy and Area under the Receiver operating curve (AUCROC) were obtained. The best ML algorithm was Gradient Boosting Tree, reaching an Accuracy and AUCROC equal to 0.968 and 0.986, respectively, when TF-IDF (Unigram) feature extraction is applied with Ten-Fold cross validation. However, the simple binary classification limited the study. Future investigation, including examination of new scenarios like multi-class classification, could confirm the potential of the proposed methodology.

Read the paper · More papers on PaperTik