Mitigating Online Abuse: Augmentation-Enhanced Deep Learning Model for Toxicity Prediction in Online Comments

Shahin Ibne Shahid · 2024

The rapid growth of online communities and social media platforms has led to a surge in toxic and harmful comments, creating a pressing need for effective automated content moderation systems. This paper presents an approach to improve toxic comment classification by employing a data augmentation method that uses paraphrasing techniques to generate diverse training samples for minority classes. By augmenting the dataset, we address class imbalance, a key issue in toxicity detection, and enhance the performance of machine learning models. We applied the augmented data to fine-tune a BERT (Bidirectional Encoder Representations from Transformers) model for toxicity detection, leveraging its deep contextual understanding of language. Our results show that the BERT model, trained on the enriched dataset, outperforms baseline models in identifying offensive language, particularly in cases where the original training data was limited or imbalanced. This method provides a scalable and effective solution to improving the accuracy and robustness of toxicity detection in online comments.

Read the paper · More papers on PaperTik