Scene-Adaptive Unsupervised Crowd Counting for Video Surveillance

Rui Ma, Yi Hou, Chenxuan Li, Huizhu Jia, Xiaodong Xie · IEEE Transactions on Circuits and Systems for Video Technology · 2025

In recent years, significant advancements in deep learning have expanded its application in a variety of computer vision tasks. However, the performance of these models heavily depends on the quality of the training data. While existing crowd counting methods yield satisfactory results on labeled datasets, they often face serious domain adaptation issues when applied to unlabeled data, the latter being more common in real-world scenarios. To mitigate this issue, we present a novel Scene-adaptive Unsupervised Crowd Counting (SUCC) framework aimed at enhancing the domain adaptability of counting models. This framework integrates a bi-branch attention network (BBA-Net) that leverages human prior knowledge to generate highly accurate density and anchor maps, which are essential for producing intermediate domain data as pseudo labels. Our SUCC framework eliminates the need for laborious manual annotation within the new data domain. Instead, it continually performs adaptive intermediate domain generation and model fine-tuning, establishing a beneficial feedback loop. Comprehensive experiments on multiple video crowd counting datasets show that our SUCC framework significantly improves domain generalizability. Furthermore, it exhibits satisfactory model stability and algorithm interpretability, attributes that are vital for the practical deployment of counting applications. The open-source code and model weights can be found on Github.

Read the paper · More papers on PaperTik