Investigation on the Clusterability of Heterogeneous Dataset by Retaining the Scale of Variables

Norin Rahayu Shamsuddin, Nor Idayu Mahat · Mathematics and Statistics · 2019

Clustering with heterogeneous variables in a dataset is no doubt a challenging process owing to different scales in a data. The paper introduced a SimMultiCorrData package in R to generate the artificial dataset for clustering. The construction of artificial dataset with various distribution helps to mimic the scenario of nature of real datasets. Our experiments shows that the clusterability of a dataset are influenced by various factors such as overlapping clusters, noise, sub-cluster, and unbalance objects within the clusters.

Read the paper · More papers on PaperTik