Explainable Ai for Offensive Content Detection and Analysis on Social Media
Alan Joseph, A K Abhinay, Anagha Tess B, Adham Saheer, Fabeela Ali Rawther, Geevarghese Titus · 2025
The widespread use of online platforms has led to an increase in inappropriate content, posing significant challenges to digital safety. This paper presents the development of a chatbot designed for offensive content analysis in online conversations. Utilizing advanced machine learning techniques, the chatbot employs a Support Vector Classifier (SVC) combined with Word2Vec for detecting offensive language. Additionally, a LIME model is introduced to enhance interpretability by analyzing the contribution of different words toward the offensiveness score. As a second phase, from the detected offensive content a Logistic Regression-based multi-label model is used to assess toxicity levels. Trained on a dataset of Twitter messages labeled for offensive and non-offensive content, the system effectively identifies and classifies harmful language. With high accuracy and computational efficiency, this solution enhances content moderation efforts and promotes a safer online environment.