Clothing Classification using Unsupervised Pre-Training
Sumeet Dhariwal, Ying Liu, Abubakrelsedik Karali, Vladimir Vlassov · 2020
Deep Learning has changed the way computer vision tasks are being solved in recent times. Deep Learning based approaches have achieved outstanding results in computer vision tasks including image classification, segmentation, object detection. Most of this success has been achieved by training deep neural networks on labelled data. In general, the more labelled data is fed to a deep learning model, the more accurate the model will be. However, labelling is time consuming and sometimes even impossible. Fashion and e-commerce are domains where a large amount of unlabelled data is available. There is a huge need to leverage these data without labels. The aim of this paper is to explore and evaluate the possibility and effectiveness of using massive amount of unlabelled data to build deep learning models. We compare the performance of these models with the performance of models built with labelled data. Specifically, we compare fully supervised deep learning with two deep learning methods with unsupervised pre-training. Our pre-trainings are based on clustering of features called DeepCluster and rotation as a self-supervision task. The comparison is performed on the DeepFashion dataset. Our experimental results have shown that using unsupervised pre-training can attain comparable classification accuracy ($\sim$1-4 % difference) on image classification comparing to fully supervised models. Furthermore, we have shown that our models uses five times less labelled data during the fine-tuning phase and still achieves comparable accuracy ($\sim$3-4 % difference) comparing to fully supervised models. These results demonstrate the potential of using unsupervised pre-training approaches in achieving comparable results to fully supervised models.