A Multilingual Intelligent System for Objectionable Text Recognition Utilizing an Explainable AI-Supported Deep Learning Model on Bengali, Bengali Transliteration, and English Embedded Text

Mohammad Sayem Chowdhury, Tofayet Sultan, Nusrat Jahan, Md. Asraf Ali, Khandaker Tabin Hasan, N Ahmed · Research Square · 2024

Abstract We live in a global society that has benefited greatly from the rise of social media, which has become a potent agent of social change. However, toxic use of text or images may damage online communities and even spark intergroup confrontations. A precise approach is needed to deal with harmful or offensive information that specifically targets people or groups. Existing literature has generally focused on a particular language of hate speech detection using typical machine learning algorithms, where a few applied deep learning, resulting in a comparatively better outcome. However, not enough work has been dedicated to Bengali or transliterating Bengali text. This means the creation of a multilingual intelligent system that is able to recognize slang and abusive language from text or photographs is the current challenge that has to be addressed for Bengali users. Based on the crisis and deep technical comparison, our research proposed the robust multilingual expert system using mBERT as a baseline model with updates utilizing global average pooling and dense dropout. Here, we received an accuracy of 92\%, higher than any of the competing methods. Then, the inclusion of OCR also evaporated the issue from the image. Additionally, we utilized LIME explainable AI to demonstrate the transparency of our model.

Read the paper · More papers on PaperTik