A Machine Learning Model to Classify Social Media Content

Naganna Chetty, Sreejith Alathur, Ashika Ruth Saldana, Amit Kr. Gupta, Mohammad Kamrul Hasan · 2024

The world has witnessed the exponential growth of internet users. The dissipation of toxic content through social media platforms has recently increased, leading to global communal activities. This work aims to develop a model to classify online toxic content. The data/tweets are collected from multiple sources and combined to train the classification model. The dataset is pre-processed to eliminate stop words, punctuations, digits, blank spaces and perform stemming. The data has been split into training and testing parts with a ratio of 75:25 respectively. The model is trained using different machine learning algorithms on extracted features from the dataset. The random forest algorithm has been shown better performance by categorizing tweets into racist, sexist, and normal than other algorithms used. The accuracy of various models is compared and tabulated. Further, the model can be enhanced by using deep learning algorithms and quality datasets.

Read the paper · More papers on PaperTik