Redefining Cyberbullying Detection with Kolmogorov-Arnold Networks(KANs) for Building Interpretable AI Beyond Black-Box MLPs

Sayan Kumar Bhowmick, Subhas Barman, Shilpa Das, Achintya Mandal, Md Shahil · 2025

Cyberbullying detection is a challenge in AI-driven content moderation, often limited by black-box deep learning models. This study adapts Kolmogorov-Arnold Networks (KAN) and its variant, FastKAN—recently proposed architectures—as interpretable alternatives to MLP classifiers, and explores one of the early applications of these models in cyberbullying detection. Our model employs a BiLSTM-based feature extractor, replacing MLP with KAN and FastKAN for transparency and efficiency.We leverage the HateXplain dataset, restructuring it to extract multiple labels, including bullying targets like race, religion, gender, sexual orientation, profession, and attacks on relatives, ensuring fine-grained bias analysis. Experiments compare MLP, KAN, and FastKAN across evaluation strategies. Results show KAN and FastKAN match or outperform MLP while enhancing interpretability and generalization, reducing bias-driven inconsistencies. Their feature extraction improves cyberbullying detection, making AI moderation clearer. This complements traditional black-box models by offering interpretable alternatives that contribute to fair and accountable AI.

Read the paper · More papers on PaperTik