Comparative Analysis of Transformer and CNN in Image Recognition Applications
Ran Mei, Shuai Tu, Kejun Chen, Ran Xiao, Yurui Yang, Yili Sun · 2024
Convolutional Neural Network (CNN) and Transformer are two mainstream model architectures for deep learning in image recognition and classification tasks. The purpose of this paper is to explore the performance differences between the two through comparative experiments, so as to provide more comprehensive guidance for model selection. In this paper, the Residual Neural Network (ResNet) and Vision Transformer (ViT) are selected. As a representative, the performance of the two in shallow and deep layers was compared from the perspectives of theoretical analysis and experimental results. The results show that the Transformer performs well in shallow networks and is suitable for handling tasks that require global feature capture. CNN, on the other hand, gradually exert their feature extraction capabilities in deep networks, which are suitable for processing complex image classification tasks. This study not only provides scientific guidance for the model selection of image recognition and classification tasks, but also helps to promote more efficient algorithm design and promote the wide application of computer vision technology in various industries.