The Application of a Hybrid Module fusing Convolutional Neural Network and Transformer in Image Classification
Tian Zhang, Deqing Zhang · 2024
In recent years, Transformers have garnered widespread attention and applications across various computer vision tasks, exhibiting significant performance improvements. However, Convolutional Neural Networks (CNNs) continue to dominate in most practical scenarios. The motivation of this study is to delve into Transformers from a practical application perspective and discover that its multi-head self-attention mechanism can effectively capture global information of features, thus endowing it with the capability for global feature modeling. Despite Transformers demonstrating remarkable capabilities in image processing applications, their computational complexity remains a significant concern. Therefore, this study aims to fuse CNNs with Transformers, constructing a novel hybrid network block called CNH-ViT, and applying it to image classification tasks. Through this fusion approach, our goal is to leverage the global capturing capability of Transformers to establish long-range dependencies between features, while using CNNs' local receptive fields to extract local texture information among features, aiming for improved performance in the field of image classification.