Explainable English Hate Speech Detection: A Custom BiLSTM Model with LIME for Multilabel Classification

Zishan Ahmed, Md. Kishor Morol, Shakib Sadat Shanto, Ahmed Shakib Reza, Md. Abdullah-Al-Jubair · 2024

The increasing prevalence of hate speech on social media has raised concerns, highlighting the need for precise and interpretable detection techniques.This research introduces a new method for identifying hate speech in English on social media.It utilizes a specifically designed Bidirectional Long Short-Term Memory (BiL-STM) model, coupled with the Local Interpretable Model-agnostic Explanations (LIME) method, to achieve this goal.This combination allows for both accurate detection of hate speech and clear explanations of why a particular piece of text is flagged as hateful.The study utilizes a combined dataset of 44,892 comments from multiple social media platforms, annotated for hate speech, offensive language, and neutral categories.Extensive experiments demonstrate the superior performance of the custom BiLSTM model compared to other machine learning and deep learning approaches, achieving an accuracy of 94.88%, precision of 94.18%, recall of 94.48%, and F1-score of 94.18%.The incorporation of LIME enhances the model's interpretability by identifying the most influential words contributing to the classification decisions.The findings highlight the potential of the proposed approach in fostering healthier online communities through accurate and explainable hate speech detection.

Read the paper · More papers on PaperTik