Scaling up the DBSCAN algorithm for clustering large spatial databases based on sampling technique

Guan Ji-hong, Zhou Shui-geng, Bian Fu-ling, Yanxiang He · Wuhan University Journal of Natural Sciences · 2001

Clustering, in data mining, is a useful technique for discovering interesting data distributions and patterns in the underlying data, and has many application fields, such as statistical data analysis, pattern recognition, image processing, and etc. We combine sampling technique with DBSCAN algorithm to cluster large spatial databases, and two sampling-based DBSCAN (SDBSCAN) algorithms are developed. One algorithm introduces sampling technique inside DBSCAN, and the other uses sampling procedure outside DBSCAN. Experimental results demonstrate that our algorithms are effective and efficient in clustering largescale spatial databases.

Read the paper · More papers on PaperTik