A Survey of Unsupervised Clustering Methods Driven by Information Completeness: Challenges, Advances, and Future Directions
Faqiang Zeng, Jimei Li, Xiangyan Tang, Yue Yang · 2024
With the increasing complexity and uncertainty of real-world data, traditional clustering methods based on complete information face challenges in dealing with incomplete data. Such methods perform poorly when faced with data sets with missing or partially incomplete data, and it is difficult to accurately reveal underlying patterns and structures. This paper investigates data clustering methods for unsupervised learning under different information completeness conditions in recent years, covering a variety of techniques from classical methods to modern deep learning architectures. We divide the existing unsupervised clustering methods into three categories based on information completeness: methods based on complete information, incomplete information, and missing information. At the same time, we also explore the challenges of these methods in practical applications, including data incompleteness, model robustness, and computational complexity. This paper also looks forward to the potential of unsupervised clustering technology in future data analysis, especially improving algorithms to deal with information missing problems to improve practical application effects.