A Comparative Study on the Impact of Intermediate DNN Features and Dimensionality Reduction in Unsupervised Data Classification Using DNN with Clustering
Jiwon Moon · 2024
The acceleration of digitalization is leading to expeditious data accumulation across all sectors of society. Additionally, as AI technology rapidly advances, the influence of AI and data is spreading across various application fields. Many AI models developed so far have been designed based on labeled data, which presents limitations when applied to the real world, where most data are unlabeled. Therefore, AI models that can effectively handle unlabeled data are becoming increasingly important. This paper aims to provide practical guidelines for deep clustering research, which combines DNN and clustering algorithms to process unlabeled data. We analyzed the impact of features extracted from specific intermediate layers of VGG16 and dimensionality reduction techniques (PCA, t-SNE) on the overall performance of deep clustering, which consists of VGG16 and KMeans. Experiments conducted using the unlabeled CIFAR-10 and MNIST datasets demonstrated that deep clustering significantly improved performance compared to KMeans clustering alone. Furthermore, deep clustering demonstrated superior performance when features were extracted from layers with an appropriate distribution of high-dimensional and low-dimensional features for clustering. It was also confirmed that applying dimensionality reduction techniques suitable for the data characteristics generally improved performance.