Analyzing Hate Speech Detection using Explainable AI
Ashish Parasar, Vandana Sharma, Preety Shoran · 2024
Hate speech in today's world especially on social media platforms are rapidly increasing rapidly and its potential to create violence, discrimination and social division. Original methods to detect hate speech such as manual moderation and keyword filtering, are becoming insufficient in managing this rapid growth of this harmful online content. This study dives into the application of XAI to enhance the transparency and interpretability of automated hate speech detection models. By incorporating XAI techniques like counterfactual explanations, feature significance analysis and saliency maps, the aim of this study is to provide the best underlying and understanding decision making process of hate speech classifiers. A comparison of various XAI methods is conducted to evaluate their effectiveness in terms of their efficiency, accuracy and interpretability. The findings of this study suggests that XAI addresses the key challenges in hate speech detection and offers valuable insights. The future research work will include development of more transparent and robust AI system for social media content filtration.