NLP-Powered Identification of Online Harassment

Kadambri Agarwal, Aditi Mittal, Anushka Gupta, Bhoomi, Sandhya Avasthi · 2024

The escalation of cyberbullying in the current digital era has become a widespread and urgent issue, endangering people’s mental and emotional health, especially young people. This study examines the field of cyberbullying detection, illuminating the approaches, tools, and moral dilemmas associated with locating and stopping this online threat. The intentional use of digital platforms to harass or hurt others is known as cyberbullying. This research presents a methodology for identifying and preventing cyberbullying, taking into consideration its hallmarks, which include repeated, purposeful harm to another person through the use of offensive and threatening words. Therefore, using the ensemble method approach, our system can identify texts that involve cyberbullying and various categories related to cyberbullying, such as age, gender, ethnicity, religion, non_cyberbullying, and other_cyberbullying. Three algorithms are combined to create our ensemble model: Random Forest, Decision Tree, and Logistic Regression. Here, to eliminate the imbalance in the data set, we additionally performed oversampling using the SMOTE approach. Additionally, after training the model, we examined the accuracy and AUC values of all models used. The most crucial thing to note is that our ensemble model has a 94.68% accuracy rate. Thus, our results demonstrated the effectiveness of our models in identifying and extenuating cyberbullying, promoting a safer online community.

Read the paper · More papers on PaperTik