Automatic Detection of Cyberbullying on Social Media Using Machine Learning
D Harish, M Alamelu, M Manimaran, Jayashakthi Vishnu P · 2023
The prevalence of digital media is expanding every day as technology advances. Individuals in the 21stcentury are growing up in an internet-equipped society. Digital media presents a lot of opportunities, but people frequently abuse it. Social media, for instance, is used to spread hatred for a person. Individuals are impacted by cyberbullying in various ways. It has an impact on more than just health; there are numerous other factors that put the victim’s life at risk. Cyber-bullying is a widespread modern event that an individual cannot completely avoid but can prevent. The method discussed in this study uses supervised-machine learning and neural networks to identify and stop cyber-bullying while considering its important attributes, such as the desire to harass a target time after time and the use of hate speech or abusive language. The models used in this study include 1-D Convolutional Neural Networks, Support Vector Machine (SVM), Logistic Regression (LR), and a Logistic Regression and Support Vector Machine (SVM) ensemble. The feature extraction methods used are Term Frequency-Inverse Document Frequency (TF-IDF) and N-grams. The optimizers used to fine-tune the models are also discussed. The method focuses on the recognition of text that is considered to be cyberbullying as well as the themes or categories that are considered to be cyberbullying, such as racism, sexual, physical mean, profanity, and many others. Results indicate that the Logistic Regression (LR) model with TF-IDF feature extraction and Stochastic Average Gradient (SAG) optimizer, outperforms existing models and methods used in the experiment, for identifying cyberbullying in text data.