A Vision Transformer with Improved LeFF and Vision Combinative Self-attention Mechanism for Waste Image Classification
Yuxiang Guo, Di Cao, Wuchao Li, Shang Hu, Jiabin Gao, Lixin Huang, Zengrong Ye · 2021
At present, many challenges exist in the application of waste image classification including high image resolution, background noise interference and low computational efficiency, which makes a great impact on the accuracy of the algorithm. Generally, convolutional neural network related strategy is adopted, which can improve accuracy to a certain extent. But it is prone to getting interfered by irrelevant information and has large size of parameters, which makes it difficult to be deployed for those resource-limited applications. Thus, a Vision Transformer is proposed. By using Vision Transformer, it significantly improves the anti-interference ability and reduces parameter scale. Furthermore, distillation strategy is adopted to improve the generalization ability of the model. In addition, this paper replaces multilayer perceptron (MLP) layer in base Vision Transformer with improved Locally-enhanced Feed-Forward (improved LeFF) to improve the convergence speed and accuracy of the model. This paper also proposes Vision combinative self-attention algorithm to reduce the requirements of computational resources especially for image with high resolution. The results show that in 79002 waste image datasets with 198 types, the accuracy of training set and test set can reach 99.98% and 95.03% respectively. Meanwhile, a relatively high accuracy rate could be achieved within less training epoch and less computational resources are required during the operation.