Investigation on the Clusterability of Heterogeneous Dataset by Retaining the Scale of Variables
Norin Rahayu Shamsuddin, Nor Idayu Mahat · Mathematics and Statistics · 2019
Clustering with heterogeneous variables in a dataset is no doubt a challenging process owing to different scales in a data. The paper introduced a SimMultiCorrData package in R to generate the artificial dataset for clustering. The construction of artificial dataset with various distribution helps to mimic the scenario of nature of real datasets. Our experiments shows that the clusterability of a dataset are influenced by various factors such as overlapping clusters, noise, sub-cluster, and unbalance objects within the clusters.