High Performance Cyberbullying Detection in Social Media: Leveraging Logistic Regression and Count Vectorizer for Enhanced Classification

Dhanroop Mal Nagar · Advances in Nonlinear Variational Inequalities · 2025

This research systematically evaluates nine supervised machine learning models, encompassing classical and ensemble approaches, for identifying cyberbullying instances within tweet-based social media content, categorizing them as 'Bully' or 'Non_bully' (Alabdulwahab et al., 2023; Ostayeva et al., 2024). Employing a rigorous methodology that includes comprehensive text preprocessing and comparison of TF-IDF and Count Vectorizer for feature extraction, along with robust evaluation via Stratified K-Fold and Stratified Shuffle Split cross-validation, the study aims to optimize classification accuracy (Al-Harigy et al., 2022; Setiawan et al., 2024). Our empirical findings, as detailed in the results section, indicate that the Logistic Regression model, paired with a Count Vectorizer and evaluated using Stratified Shuffle Split, achieved the highest F1-Score of 0.8460, an accuracy of 0.8467, and a recall of 0.9067. This performance highlights the effectiveness of simpler linear models when combined with raw term frequency features and robust cross-validation in cyberbullying detection (León-Paredes et al., 2023; Muneer & Fati, 2020; Salawu et al., 2017). The Count Vectorizer consistently demonstrated superior discriminative power compared to TF-IDF for this dataset, suggesting that raw term frequencies were more indicative of cyberbullying patterns (Talpur & O’Sullivan, 2020). This work contributes to the academic understanding of cyberbullying detection by systematically identifying optimal model and feature engineering strategies, offering practical insights for developing effective automated systems to combat online harassment (Raj et al., 2021)

Read the paper · More papers on PaperTik