Enhancing Image Clustering with CLIP
Mengjuan Li, Wenming Cao, Lei Zhang, Man Li, Mingming Yang · 2024
Both similarity of instances and difference among clusters play an essential role in improving the performance of deep image clustering. Existing deep clustering methods often fail to capture both consistency and distinctness of semantics when the label information is not available. In this work, we propose a semantic prototype pseudo-label-based image clustering model, which uses a CLIP image encoder as the feature model and incorporates an ensemble-based clustering head. The feature model is responsible for evaluating the consistency between instances, while the clustering prediction head is used to discern semantic differences. The prototype pseudo-label algorithm is designed to train the clustering head to effectively improve the accuracy and reliability of self-supervision. The optimization of the clustering network is divided into two stages, both of which do not rely on real labels: 1) The feature model is trained using contrastive loss to evaluate the consistency of instances; 2) The clustering head is trained using the prototype pseudo-label algorithm to distinguish between semantic categories. Comprehensive experimental evaluations demonstrate the effectiveness of our method in improving clustering performance, especially achieving a 10% improvement on three recognized evaluation metrics compared to existing techniques on the STL-10 and ImageNet-10 datasets.