Integrating BERT and RESNET50V2 for Multimodal Cyberbullying Detection*
Immaculate Musyoka, John Wandeto, Benson Kituku · 2024
Despite a range of advantages that social networks provide, cyberbullying is widespread. Several tools have been developed to automatically detect text-based cyberbullying, but with increased use of a combination of image and text in social media content, there is a need to detect cyberbullying from multimodal content such as images and text. In this paper, we present a custom architecture that uses two pretrained models for feature extraction, BERT and RESNET50V2. The proposed custom model architecture was trained on 149,823 instances of image-text pairs and produced a testing accuracy of 78%. Cross attention and use of dropouts were employed during model training, and focal loss was also used to make the model perform better on the rare class. A separate model that used state-of-the-art BERT for both feature extraction and classification of text only was trained on 149,823 instances of text and achieved an accuracy of 75%. The custom image and text model produced better results in comparison to state-of-the-art BERT for text, with its evaluation showing a good capability of detecting more cyberbullying in multimodal data.