Large Language Model Used to Improve the Ability of Detecting Cyberbullying
Haoyuan Wang · Applied and Computational Engineering · 2025
For cyberbullying detection in online platforms, balancing computational efficiency and contextual accuracy remains a critical challenge. Traditional machine learning (ML) models often lack deep semantic understanding, while standalone large language models (LLMs) suffer from prohibitive latency in real-time applications. To tackle this, we propose a two-stage hybrid framework: Stage 1 employs SVM or Random Forest for rapid filtering, and Stage 2 uses a lightweight DistilBERT for contextual analysis. Evaluated on a 50,000-sample English dataset from the Cyberbullying Research Center, comprising 40,000 training, 5,000 validation, and 5,000 test samples with a class distribution of 5% bullying (2,500 samples) and 95% non-bullying—the framework achieves an F1-score of 0.89, outperforming traditional ML (0.80) and standalone LLMs (0.87), while delivering 83% faster inference than full LLM processing. Preprocessing involved lowercasing, SpaCy tokenization, and NLTK stopword removal. Key strengths include real-time scalability and enhanced detection of implicit bullying. However, limitations persist: dependency on Stage 1 filtering accuracy and challenges in interpreting sarcasm (contributing to 22% of errors) and novel slang, which necessitate further research.