Hierarchical Fusion of 3D CNNs with Confidence Awareness for Violence Recognition in Videos

Nadjia Khatir, Hassina Meziane · Complex Systems Informatics and Modeling Quarterly · 2025

The deployment of surveillance networks in smart cities plays a pivotal role in enhancing public safety through the monitoring of various environments such as roads, airports, residential areas, and establishments. Nevertheless, the vast volumes of video data generated daily by these networks present both opportunities and challenges in terms of information management and analytical processing. In this study, we propose a novel trust-aware fusion framework of video-based violence and threat modeling by combining two state-of-the-art models. I3D, which excels in overall spatio-temporal reasoning, and C3D, which learns short-term motion behaviors. In Stackelberg’s game theory, the process of fusion outlines inference as a sequential decision-making process, wherein the leader is I3D, and C3D acts as a follower. A dynamic confidence threshold governs the prediction delegation power, enabling adaptive decision-making based on model confidence. Extensive experiments on a three-class dataset (Normal, Violence, Weaponized) prove that the introduced fusion strategy significantly outperforms single models. Setting the confidence threshold to 0.5 achieves 97.27% peak of overall accuracy. In addition, class-wise performance reveals considerable improvements, especially in the Violence class, where precision is 99% and the F1 score is 94%, versus 82% and 85% when using I3D individually. The experiments confirm the performance of the confidence-aware fusion for robust and context-adapted threat detection in smart-city surveillance.

Read the paper · More papers on PaperTik