S3T: A New Self-Supervised Learning with Swin Transformer
Hirak Mazumdar, Sriparna Saha, P Bhargav, MSVPJ Satvik · 2024
This paper conducts a detailed examination of the Swin Transformer’s utility as an encoder within the context of the Barlow Twins method, with a focus on its application to the CIFAR-10 dataset. The growing interest in self-supervised learning methods stems from their capacity to derive meaningful representations from unlabeled data. The Barlow Twins method, specifically, exploits two augmented perspectives of the same image, emphasizing the maximization of cross-covariance between these perspectives to acquire robust representations. In our pursuit of enhancing the method’s efficacy, we opted to replace the conventional ResNet-18 encoder with the Swin Transformer—a transformative, transformer-based architecture that has recently exhibited outstanding performance across various image recognition tasks. The success achieved by the Swin Transformer within the Barlow Twins framework underscores its adaptability and versatility for applications extending beyond image recognition. Much like how transformers have revolutionized natural language processing, integrating them into self-supervised learning approaches holds the potential to catalyze transformative advancements in computer vision tasks. Through our experimentation on the CIFAR-10 dataset, we establish that the Swin Transformer’s attention mechanism adeptly captures intricate patterns and relationships within images, signaling its capacity to revolutionize a spectrum of image-related applications. Moreover, our exploration of Swin Transformers in the realm of self-supervised learning opens up intriguing avenues for further research. This includes harnessing these architectures for transfer learning, semi-supervised learning, and other real-world scenarios where acquiring labeled data proves challenging or expensive. As the landscape of self-supervised learning continues to evolve, the adaptability and efficacy of the Swin Transformer provide a promising foundation for future research and innovation within the realms of machine learning and computer vision.