Multi-Model Ensemble Approach for Enhanced Cyberbullying Detection Across Diverse Categories

Saraswati Patil, Shubhan S. Punde, Prince Sahani, Abhinav Salve · 2024

This research study presents an advanced multi-model ensemble approach for cyberbullying detection that effectively identifies and classifies various forms of online harassment across social media platforms. The proposed framework combines five machine learning models - LightGBM, XGBoost, Support Vector Machine (SVM), Random Forest, and Logistic Regression - through a voting classifier mechanism. Our methodology incorporates sophisticated text preprocessing techniques and leverages both TF-IDF and Word2Vec embeddings for feature extraction, while addressing dataset imbalances through the Synthetic Minority Over-sampling Technique (SMOTE). The ensemble approach achieves an impressive accuracy of approximately 94%, outperforming individual models which scored between 92-93%. Particularly notable are the model's strong performance metrics across different cyberbullying categories, with F1-scores ranging from 0.84 to 0.99 for detecting gender-based, religious, age-related, and ethnic harassment. The research also introduces a user-friendly Streamlit interface for real-time cyberbullying detection, making the technology accessible for practical applications. Our findings demonstrate that combining multiple well-tuned models with advanced feature engineering techniques can significantly enhance the accuracy and reliability of cyberbullying detection systems.

Read the paper · More papers on PaperTik