SAFEPLAY-X: A comprehensive gameplay video dataset for violence detection with explainable deep learning applications
Aakash Singh, Vinayak Bansal, Muskan Saini, Deepawali Sharma, Vivek Kumar Singh · Expert Systems with Applications · 2026
In recent years, gameplay videos have grown rapidly in popularity on platforms like YouTube, attracting large audiences, including children and adolescents. These videos often offer entertainment and engagement, but there is growing concern about the kind of content they contain. In particular, some of these videos feature intense scenes of violence, which may not be suitable for younger audiences and could have a negative impact on their mental well-being. To address this issue, a novel dataset called SAFEPLAY-X is introduced which is annotated by three independent annotators. The dataset consists of 37,299 five-second gameplay video clips collected from YouTube to detect violent content. It supports two subtasks: binary classification of violent vs. non-violent videos, and multi-class classification of violent content into three categories-threat of violence, physical assault, and lethal violence. Both unimodal models for audio and video, and multimodal models are implemented on this dataset. Our proposed multimodal model, which combines MC3_18 model for video and ResNet model for audio, outperforms the unimodal models, achieving 95.75% accuracy for binary classification and 81.91% for multi-class classification. Additionally, Grad-CAM visualizations for both modalities is used to improve the interpretability and transparency of the model’s predictions. SAFEPLAY-X is the first dataset of its kind focused on violence detection in gameplay videos.