Where and What We Should Look at: A Cross-Modal Fusion Method with Attention for Object Counting
Shiyu Zhao, Donglai Wei, Liang Song · 2022
Object counting is a challenging computer vision task. Recent studies find that incorporating multi-modal information can help find the hard points. But the cross-modal representation learning, which aims to comprehend and represent the information in various modalities, is still far from being fully explored in this area. The current works fail to model the dynamic complementary relationship between different modalities, leading to incomplete information extraction. To make up for that, we propose an attention mechanism to facilitate information fusion. Our work adopts a three-stream structure as the framework, consisting of two modality-specific branches and one modality-shared branch. The proposed two attention modules aggregate contextual statistics along the channel and spatial dimensions to generate attention maps that reflect the features' importance in various modalities. In this way, the complementary relationship between different modalities would be exposed to our model precisely. Besides, we share the convolution filters of the attentions modules between the RGB branch and modality-shared branch. The modality-common statistic can be captured through parameter sharing, exposing the informative features. To validate our ideas, we implement extensive experiments on the RGBT -CC benchmark. Consequently, we outperform the state-of-the-art results by a large margin. The performance strongly supports the effectiveness of our method.