SC-ViT: Semantic Contrast Vision Transformer for Scene Recognition
Jiahui Niu, Xin Ma, Rui Li · 2024
Scene recognition remains a challenging task in image recognition. Despite the remarkable advances made by deep learning, especially with the emergence of Convolutional Neural Networks (CNNs), scene recognition continues to face unresolved issues. Due to high complexity of scene images, merely identifying several objects in the image is insufficient for obtaining accurate results. Furthermore, the wide range of scene categories makes single-modality learning susceptible to confusion. To address these challenges, we propose an end-to-end multimodal network, SC-ViT, based on the Vision Transformer (ViT). Leveraging the powerful self-attention mechanism, our model captures visual and contextual cues from both RGB images and semantic information. The semantic information is derived through semantic segmentation, encompassing object category information and spatial layout within the scene. Combining semantic information with RGB images results in a comprehensive representation of the scene. Specifically, we utilize two branches, each equipped with self-attention mechanisms but with different structures, to extract features from RGB images and semantic information. Through a contrastive learning framework, SC-ViT aligns the feature representations of RGB and semantic modalities, enhancing their consistency and discriminative power to express the scene. Experimental evaluations on the MIT Indoor67 and SUN397 datasets demonstrate that SC-ViT outperforms state-of-the-art methods, achieving significant improvements in scene recognition.