Multimodal Detection of Offensive Content in Hindi Memes
Kriti Dubey, Vaishnavi Srivastava, Garima Sharma, Nonita Sharma, Deepak Kumar Sharma, Uttam Ghosh, Osama Alfarraj, Amr Kamal Rabea Tolba · ACM Transactions on Asian and Low-Resource Language Information Processing · 2025
Activities like sharing of thoughts, advertising business, connecting with peers, and staying updated in this world are highly facilitated by social media platforms. In the unique form of media called memes, information is conveyed through the image-to-text or text-to-image dependency relationship. Popular memes are often driven by viewers rather than marketing or advertising techniques, which indicates the level of engagement of social media users with memes. Given the popularity of memes, a demand for a solution to identify and counteract hate-spreading memes on social media platforms is raised. In this study, a multimodal machine learning approach to detect offensive memes is presented, where the text of memes are spelled in the Devanagari script of the Hindi language. A dataset of 9262 images has been created, and they have been labeled as offensive or not offensive. As the dataset is highly imbalanced, another dataset which has 3732 images is created by undersampling the total dataset and the models are trained for both the imbalanced dataset as well as the balanced dataset. Finally, the classification problem is solved through a multimodal Logistic Regression classifier that utilizes concatenated feature representations of image and text. An accuracy of 0.81 is achieved by the model.