A Review of CNN and Transformer Applications in Image Processing

Zhu Cheng, Niu Xiaoqin · 2024

Convolutional Neural Network (CNN) and vision Transformer are two important deep learning models in the field of image processing, and they have made remarkable achievements in this field after years of continuous research and progress. In recent years, the hybrid model of CNN and vision Transformer is gradually emerging. Extensive research has constantly overcome the weaknesses of the two models, and effectively play their respective highlights, showing excellent results in image processing tasks. This paper is based on the hybrid model of CNN and vision Transformer. First, the architecture, advantages and disadvantages of CNN and Vision Transformer model are summarized, and the concept and advantages of hybrid model are summarized. Secondly, the classification and main representative models of hybrid models are reviewed, and their applications in specific image processing tasks are described from multiple perspectives. Finally, the future research direction of hybrid model is deeply analyzed, and the prospect is put forward.

Read the paper · More papers on PaperTik