MVT: Using Mixup in Vision Transformers for Enhanced Data Augmentation

Tanmay Jain, Vishal Yadav, Tanishq Kinra, Shailender Kumar · 2023

Using large deep neural networks for image classification can be very effective, but it has some negative effects such as memorization and vulnerability to "adversarial examples". To tackle such issues we propose to apply mixup [4] in various parts of attention space of vision transformer [1] architecture. Mixup is a data augmentation technique which involves training a neural network on pairs of examples and their associated labels with pre defined weights associated to each input. By doing so, the network is encouraged to behave simply linearly between training instances, thereby acting as a form of regularization. While previous work on mixup has focused on the image domain, we introduce it here in the context of the encoder blocks of the vision transformer and explore its impact on classification accuracy across different datasets.

Read the paper · More papers on PaperTik