The problem of high dimensionality with low density in clustering
T. Sudha, Swapna Sree Reddy. Obili · International Journal of Managment, IT and Engineering · 2012
In many real-world applications, there are a number of dimensions having large variations in a dataset. The dimensions of the large variations scatter the cluster and confuse the distance between two samples in a dataset. This degrades the performances of many existing algorithms. This problem can be happened even when the number of dimensions of a dataset is small. Moreover, no existing method can distinguish whether the dataset has the highly repeated problem or low-density's problem. The only way to distinguish the problem is by a prior knowledge, which is given by the user.