A preprocessing data-driven pipeline for estimating number of clusters
Michal Koren, Or Peretz, Oded Koren · Engineering Applications of Artificial Intelligence · 2024
Due to the abundance of information, Artificial Intelligence (AI) research and development is a major focus in academia and industry. As the data volume increases, organizations face challenges in defining, selecting, and prioritizing the most relevant and influential variables for research, growth, and development. In this work, a preprocessing and data-driven pipeline will be demonstrated and implemented in estimating the number of clusters for a given dataset. The pipeline includes two main procedures: 1) a threshold-based method for feature selection, and 2) an estimation of neighborhood size (i.e., the total number of clusters). The procedure uses a risk parameter representing the number of features the user is willing to risk lowering in the overall correlation. This study is significant in the proposed pipeline's ability to predict the number of clusters of a given dataset by controlling the feature correlation and threshold learning of clustering evaluation measures. Using a data-driven methodology, the pipeline estimates the number of desired clusters for a given dataset to understand data distribution before evaluating any clustering procedure. The results showed the pipeline method achieved an improvement of up to 157.805% in the silhouette coefficient. In addition, there were also consistent improvements in other clustering performance metrics.