Learn Feature Representation from Unlabeled Data using Self-Supervised Learning

Avani Khokhariya, Amit R Thakkar, Nikita Bhatt, Rohan Vaghela · 2024

In the recent past, promising performance is achieved by supervised deep learning networks in various domains like computer vision, machine translation, speech recognition, natural language processing (NLP), etc. But current deep-learning approaches required a large amount of manual data to achieve promising performance. However, manual annotation of data is time-consuming and needs human expertise, which is always not possible in a domain like healthcare. Recently an alternative method called Self Supervised Learning (SSL), uses unlabeled data and does not require manual annotation. With the help of a pretext task, SSL extracts and learns features from the unlabeled data, enabling models trained for these tasks to acquire latent representations that enhance subsequent tasks like object detection and classification. In this work, the unlabeled data is fed to the convolution neural network (CNN), which learns features and transfers them to the downstream task to generate the labeled data. The experiments are conducted on the fashion MNIST dataset in a self-supervised learning environment and a comparison is made with a different size of training data and the rotation pretext task approach is compared with SimCLR approach of contrastive learning which achieve state-of-art performance. Another experiment performs on Chest X-ray dataset, COVID-19 CT Scan dataset using SimCLR and other methods of contrastive learning. The experiment concludes that contrastive learning of self-supervised learning achieves better performance over rotation pretext task of self-supervised learning. The self-supervised learning environment is preferable in various domains where human annotation for labeling datasets is very costly or not feasible.

Read the paper · More papers on PaperTik