Joint Contextual Transformer and Multi-scale Information Shared Network for Crowd Counting
Xin Zeng, Shizhe Hu, Huake Wang, Jinna Zhang · 2022 5th International Conference on Pattern Recognition and Artificial Intelligence (PRAI) · 2022
Crowd counting is a fundamental task in computer vision with applications for video surveillance and public safety management. However, scale variation remains a challenge. To overcome this weakness, we propose a contextual transformer and multi-scale information shared network (CTrans-MISN) to improve the performance of the counting. It consists of (i) a contextual transformer, capturing the contextual information of crowds, that makes the network more powerful to face the scale variation. And (ii) with the help of multi-scale information shared module, the model can learn more conducive information to identify congested crowd scenes accurately. We finally assess the proposed method on four public counting datasets. The extensive experimental results demonstrate the superiority of our approach over other compared baselines.