Can Unlabelled Data Improve AI Applications? A Comparative Study on Self-Supervised Learning in Computer Vision.

Markus Bauer, Christoph Augenstein · Annals of Computer Science and Information Systems · 2023

Artificial Intelligence (AI) represents a highly investigated area of study at present and has already become an indispensable component within an extensive range of business models and applications.One major downside of current supervised AI approaches lies in the need of numerous annotated data points to train the models.Self-supervised learning (SSL) circumvents the need for annotation, by creating supervision signals such as labels from the data itself, rather than requiring experts for this task.Current approaches mainly include the use of generative methods such as autoencoders and joint embedding architectures to fulfil this task.Recent works present comparable results to supervised learning in downstream scenarios such as classification after SSL-pretraining.To achieve this, typically modifications are required to suit the approach for the exact downstream task.Yet, current review works haven't paid too much attention to the practical implications of using SSL.Thus, we investigated and implemented popular SSL approaches, suitable for downstream tasks such as classification, from an initial collection of more than 400 papers.We evaluate a selection of these approaches under real-world dataset conditions, and in direct comparison to the supervised learning scenario.We discuss SSL's potential to take up with supervised learning, as well as the influence of the right training methods.Furthermore, we also introduce future directions for SSL research, as well as current limitations in real-world applications.

Read the paper · More papers on PaperTik