An Efficient Clustering Approach using DBSCAN

R. V. Sravana Kumar · Helix · 2018

With the fast development of the Internet, the information available in the database showing explosive growth where it wouldn't be that easy to obtain desired information accurately from the available data.This unexpected growth of internet has also led to the frequent usage of e-commerce sites which concentrate on sales based on their relevant products.To do so clustering of products is done so as to identify the sales with profits and which are lagging behind by finding similarities among the products.The clustering analysis is done with the help of various clustering algorithms like DBSCAN, K-means, GRID based, Hierarchical based, etc., .Each algorithm will produce different results; you'll never be certain whether one result is better than the other or even whether the result is of any value.Kmeans algorithm is preferable choice for datasets having small number of clusters with proportional sizes and linearly separable data and also you can scale it to use on very large data sets.K-means algorithm needs the number of clusters fixed in advance and it arises problems when the data contains outliers.Also the downside of K-means is it leads to problems when clusters are of different sizes, densities and non-globular shapes.The DBSCAN implementation does not require any user-defined initialization parameters to create an instance like Kmeans do.So DBSCAN algorithm is more advantageous when compared to other algorithms.It is because DBSCAN doesn't need one to specify the amount of clusters within the knowledge a priori, as against K-means.DBSCAN will notice different arbitrary formed clusters.Density clustering (DBSCAN) seems to correspond more to human intuitions of clustering, rather than distance from a central clustering point (K-means).

Read the paper · More papers on PaperTik