Clustering and visualization of a high-dimensional diabetes dataset
Piotr Lasek, Zhen Mei · Procedia Computer Science · 2019
Data clustering algorithms have proved to be important and widely used methods of artificial intelligence and data mining for discovering unknown yet important patterns in datasets. Nevertheless, one of the additional aspects of data clustering is proper interpretation of the clustering results. In this paper we aim to investigate possibilities of using both data clustering and visualization methods to analyze a sample diabetes dataset. In the first part, we focus on how to cluster a highly-dimensional sample dataset and then, we concentrate on how to properly visually present the clustering results in the most meaningful way to uncover potentially interesting behavioral patterns or features of diabetes patients. In this work we examine two clustering algorithms (DBSCAN, k-Means) along with several different distance measures. We also present sample visualizations of clustering results generated by an application which we have developed and discuss if the proposed way of clustering results visualization can be helpful in understanding the analyzed dataset and lead a viewer to drawing valuable conclusions about it.