Unified Transformer Framework for Automated Cyberbullying Detection
Enas Ahmad Alikhashashneh, Hedaia Alsawan, Khalid M.O. Nahar, Nahlah M. Shatnawi, Ammar Almomani, Mohammad Alauthman, Shavi Bansal, Vincent Shin-Hung Pan · International Journal of Cloud Applications and Computing · 2025
Cyberbullying is a fast-growing public-health hazard, demanding reliable, real-time detection of abusive language online. This study presents a unified transformer framework that compares bidirectional encoder representations from transformers, generative pre-trained transformer-2 and text-to-text transfer transformer (T5) on the 90 356-message Mendeley Cyber-Bullying corpus. A shared pipeline normalises text, removes stop-words, and using T5, augments minority classes to curb imbalance. Models are fine-tuned under identical splits (70% train/15% val/15% test, 15 epochs) and scored with accuracy, precision, recall, and F1. Augmented T5 leads with 92.7% accuracy, surpassing generative pre-trained transformer-2 (90.1%) and bidirectional encoder representations from transformers (89.4%). Confusion-matrix analysis shows T5 best balances true- and false-positive rates. Results validate (a) casting cyberbullying detection as sequence-to-sequence; (b) transformer-driven augmentation as an efficient remedy for skewed data; and (c) the feasibility of lightweight, fine-tuned transformers for scalable safety tool.