A Knowledge-Enhanced Approach for Robust Online Harassment Detection using XGBoost within Gradient Boosting
J. Sathya, F. Mary Harin Fernandez · 2023
Digital communication platforms have increased online harassment, requiring sophisticated methods for detecting and preventing it. This study introduces a comprehensive strategy for spotting online harassment, utilizing a specialized knowledge framework, advanced feature extraction methods, and the robust XG Boost algorithm within the Gradient Boosting framework. To identify online harassment tendencies, the proposed approach combined TF- IDF, n-grams, and word embeddings in phase two of feature extraction. The core of proposed method hinges on implementing the XG Boost algorithm within the Gradient Boosting paradigm. Empirical assessments performed on a varied and representative dataset demonstrate the effectiveness of the proposed strategy. Metrics like accuracy, precision, recall, F1-score, and ROC-AUC collectively highlight its resilience in differentiating instances of online harassment from harmless content. Furthermore, The Proposed model (XGBOOST) exhibits the highest accuracy at 92%, outperforming other models such as LIGHTLGM (90%), CATBOOST (88%), and N-GRAMS (84%). Additionally, XGBOOST excels in the precision, recall, F1-score, and MCC XG Boost's prowess in Gradient Boosting and the integration of domain-specific knowledge frameworks contributes to a comprehensive remedy for detecting online harassment.