Overcoming Domain Shift in Violence Detection with Contrastive Consistency Learning

Zhenche Xia, Zhenhua Tan, Bin Zhang · Big Data and Cognitive Computing · 2025

Automated violence detection in video surveillance is critical for public safety; however, existing methods frequently suffer notable performance degradation across diverse real-world scenarios due to domain shift. Substantial distributional discrepancies between source training data and target environments severely hinder model generalization, limiting practical deployment. To overcome this, we propose CoMT-VD, a new contrastive Mean Teacher-based violence detection model, engineered for enhanced adaptability in unseen target domains. CoMT-VD innovatively integrates a Mean Teacher architecture to adequately leverage unlabeled target domain data, fostering stable, domain-invariant feature representations by enforcing consistency regularization between student and teacher networks, crucial for bridging the domain gap. Furthermore, to mitigate supervisory noise from pseudo-labels and refine the feature space, CoMT-VD incorporates a dual-strategy contrastive learning module. DCL systematically refines features through intra-sample consistency, minimizing latent space distances for compact representations, and inter-sample consistency, maximizing feature dissimilarity across distinct categories to sharpen decision boundaries. This dual regularization purifies the learned feature space, boosting discriminativeness while mitigating noisy pseudo-labels. Broad evaluations on five benchmark datasets unequivocally demonstrate that CoMT-VD achieves the superior generalization performance (in the four integrated scenarios from five benchmark datasets, the improvements were 5.0∼12.0%, 6.0∼12.5%, 5.0∼11.2%, 5.0∼11.2%, and 6.3∼12.3%, respectively), marking a notable advancement towards robust and reliable real-world violence detection systems.

Read the paper · More papers on PaperTik