Multi-constraints active learning assisted deep-ensemble spatio-textural feature learning model for violence detection in surveillance dataset

Duba Sriveni, R. Loganathan · Connection Science · 2025

This paper presents Vi-mCALNET, a multi-constraint active learning-assisted deep-ensemble spatio-textural feature learning model for violence detection in surveillance videos. Unlike traditional approaches, Vi-mCALNET integrates active learning-driven frame selection with deep ensemble learning to enhance classification accuracy while reducing computational complexity. Traditional deep learning vision models for violent crime detection face limitations including inability to use contextual details, long-term dependency issues, gradient vanishing, and accuracy degradation. Vi-mCALNET addresses these challenges through a comprehensive approach. The model employs GLCM, ResNet101, and DenseNet121 for feature extraction, followed by a heterogeneous ensemble classifier comprising SVM, DT, k-NN, NB, and RF. Extracted features are fused into a composite feature vector, processed through PCA and z-score normalization to prevent local minima, convergence issues, and overfitting.The heterogeneous ensemble classifier uses maximum voting to classify videos as violent or non-violent. Vi-mCALNET achieved superior performance with 99.51% accuracy, 99.32% precision, 99.36% recall, and 0.994 F-measure on publicly available datasets.Ablation studies and statistical significance analysis confirmed Vi-mCALNET's robust performance with lower variance, making it suitable for real-time, scalable surveillance applications while reducing annotation costs and computational demands.

Read the paper · More papers on PaperTik