Offensive Text Detection on Social Media Using Machine Learning

Jesslyn Tanuwijaya, Ivan Sebastian Edbert, Hans Christian Arinardi, Derwin Suhartono · 2024

People are social creatures, people of all ages use social media to engage with one another. But not everything that circulates on social media is suitable for all ages; some content, such hate speech, bullying, and cursing, can be harmful to users' mental and physical health. Thus, the necessity for a social media content filtering tool arises. To solve this problem, the researcher proposed a model to classify English tweets by implementing machine learning algorithms specifically Random Forest Classifier, SVM, XGboost, Logistic Regression were used to detect and classify offensive text messages. In this case, dataset were extracted by using countvectorizer algorithm, the researchers also apply data cleaning to remove stop words, punctuation, url and prefix that come from the dataset. Because the imbalance of the dataset the researchers try to fix by using undersampling technique. The best result of the experiments are Random Forest Classifier perform 0.8808 accuracy.

Read the paper · More papers on PaperTik