Learning with Balanced Criss-Cross Attention for Cross-Modality Crowd Counting
Xin Zeng, Wanjun Zhang, Huake Wang, Xiaoli Bian · 2023
Cross-modality crowd counting is one of the most essential tasks in multimedia and image processing, which usually uses multi-sensor information as input in neural networks. Various approaches have been proposed to extract the alignment and relationships between the different modalities in the task of crowd counting. In this work, we explore how to further remedy the cross-modal discrepancies and learn latent relevance across different modalities. We present a novel RGBT crowd counting framework, namely Balanced Criss-Cross Attention Network (BCANet), to overcome the above limitations. To bridge the two modalities, we introduce a Balanced Criss-Cross Attention (BCA) module to encode complementary information across modalities. Lastly, we evaluate our BCANet via extensive experiments and demonstrate that it consistently achieves state-of-the-art results on RGBT-CC and DroneRGBT datasets.