INNOVATIVE HYBRID TRANSFORMER MODEL FOR INTELLIGENT VIOLENCE RECOGNITION IN SURVEILLANCE SYSTEMS
Sumeet Kothari · International Journal of Apllied Mathematics · 2025
Ensuring public safety in today’s rapidly expanding urban environments requires intelligent and automated surveillance systems capable of identifying violent incidents in real time. Traditional deep learning approaches such as CNNs, MobileNet, YOLO, and Ensemble models have shown promise but remain limited in capturing long-range temporal dependencies and global contextual information, resulting in lower accuracy, instability, and weaker generalization across diverse datasets. To address these shortcomings, this study introduces an Innovative Hybrid Transformer Model for Intelligent Violence Recognition in Surveillance Systems, which integrates spatial feature extraction through multi-scale CNNs, temporal sequence learning via BiLSTM, and cross-attention Transformers to model global spatio-temporal relationships. The framework was evaluated on benchmark datasets, including the Violence Dataset and the Road-Anomaly Dataset, achieving superior performance with ≈99% accuracy, 99% precision, 99% recall, and minimal error rates, significantly outperforming state-of-the-art CNN and YOLO architectures. Furthermore, optimization techniques such as pruning reduced parameters to below 2M, ensuring real-time applicability with inference speeds of under 100 ms per frame, making it feasible for deployment in smart city environments. The proposed model not only demonstrates robustness and scalability but also lays the groundwork for more context-aware, reliable, and efficient surveillance systems. Future work will focus on incorporating multimodal data, including audio and sensor inputs, and enhancing explainability to ensure trustworthy deployment in real-world public safety applications.