Robust Tokenizer for Vision Transformer
Rinka Kiriyama, Akio Sashima, Ikuko Shimizu · 2023
Vision Transformer is the State of the Art(SOTA) method in the image processing domain, but its robustness against adversarial examples still needs to be improved. This paper proposes Tokenizer called "ArbViT" (Arbitrary patches for Visual Transformer). To enhance the robustness to adversarial attacks, we introduce our Tokenizer for feature processing before Transformer. Specifically, when converting image patches into embedded features (or tokens), we do not use weights but reorder the features. In experiments, we show the results by introducing ArbViT and by the conventional method, ViT. Although no clear difference in accuracy between the model using ArbViT and the original ViT, we show that the generated model by introducing ArbViT is more robust against the adversarial attacks than the method using the conventional Tokenizer.