An Optimised BERT Pretraining Approach for Identification of Targeted Offensive Language: Data Imbalance and Potential Solutions

Ruth Mifsud, Lipika Deka, Indrani Lahiri · 2023

Targeted offensive comments and hate speech on online media platforms are on the rise, with evidential mental health consequences including suicide. Several NLP techniques have been proposed and in use. However, data imbalance in the training dataset is stopping them from performing at full potential. Solutions include under-sampling of the majority class, oversampling of the minority class or introducing synthetic samples. These approaches present with their own unique problems - that of critical information loss, overfitting and non-generalised models. The presented research explores these approaches for addressing the data imbalance problem, by varying the under/over/synthetic sampling rate and studying the performance as well as the generalisability of the models.

Read the paper · More papers on PaperTik