Adaptive Methods for Determining DBSCAN Parameters

Kedar Sawant · 2014

Emergence of modern techniques for scientific data collection has resulted in large scale accumulation of data pertaining to diverse fields. Cluster analysis is a primary method for database mining [8]. Among different types of cluster the density cluster has advantages as its clusters are easy to understand and it does not limit itself to shapes of clusters. But existing density-based algorithms are lagging behind. Almost all of the well-known clustering algorithms require input parameters which are hard to determine but have a significant influence on the clustering result. Furthermore, for many real-data sets there does not even exist a global parameter setting for which the result of the clustering algorithm describes the intrinsic clustering structure accurately [1][2]. This paper gives a survey of density based clustering algorithms. DBSCAN [15] is a base algorithm for density based clustering techniques. It can detect the clusters of different shapes and sizes from large amount of data which contains noise and outliers. The main drawback of traditional clustering algorithm was largely recovered by VDBSCAN algorithm. But in VDBSCAN algorithm the value of parameter ‘K’ was a user input dependent parameter. It largely degrades the efficiency of permanent Eps. In our proposed method the Eps is determined by the value of ‘k’ in varied density based spatial cluster analysis by declaring ‘k’ as variable one by using algorithmic average determination and distance measurement by Cartesian method and Cartesian product on two dimensional spatial dataset where data are sparsely distributed. So the objective is to enhance the existing DBSCAN algorithm by automatically selecting the input parameters and to find the density varied clusters. The proposed algorithm discovers arbitrary shaped clusters, requires no input parameters and uses the same definitions of DBSCAN algorithm.

Read the paper · More papers on PaperTik