Addressing Data Imbalance in Machine Learning Solutions for Sexual Violence Detection

Azani Cempaka Sari, Karto Iskandar, Amrita Prasad, Risma Yulistiani, Aditya Fahrizal Kurniawan, Nora Fitriawati, Setiawan Joddy, Rico Halim · 2024

The case of sexual violence has become rampant especially with the introduction of the internet, social. Platforms act both as a means of perpetuating this violence through harassment of victims and as a platform for fighting back by sharing incidences of such occurrences. It is immediately crucial to obtain highly effective and accurate classification models for early identification and control of this behavior. This research aims to address data imbalance in detecting text-based sexual harassment in Indonesia using machine learning techniques by finding better ways of mitigating the imbalance that exists and where non-harassment cases dominate the dataset rather than harassment cases. This study is based on data extracted from Instagram, TikTok, and X (previously Twitter) with the Synthetic Minority Over-sampling Technique (SMOTE) to overcome the imbalance of data. Specifically, three machine learning algorithms were used: Support Vector Classifier (SVC), Random Forest Classifier, and Naïve Bayes Classifier were used. The SVC model had the highest accuracy of 97 percent under the accuracy, precision, recall, and F1-score categories. The findings clearly suggest that when applied with SMOTE as a preprocessing and SVC classifier, the problem of data imbalance is solved and harassment is detected more accurately. This paper further advances the area of Artificial Intelligence in online harassment, providing a novel approach to reach safer internet space and improve the assistance to the victims.

Read the paper · More papers on PaperTik