GF-DBSCAN: a new efficient and effective data clustering technique for large databases

Cheng-Fa Tsai, Chien-Tsung Wu · 2009

The DBSCAN data clustering accurately searches adjacent area with similar density of data, and effectively filters noise, making it very valuable in data mining. However, DBSCAN needs to compare all data in each object, making it very time-consuming. This work presents a new clustering method called GF-DBSCAN, which is based on a well-known existing approach named FDBSCAN. The new algorithm is grid-based to reduce the number of searches, and redefines the cluster cohesion merging, giving it high data clustering efficiency. This implementation has very high clustering accuracy and filtering rates, and is faster than the famous DBSCAN, IDBSCAN and FDBSCAN schemes. Key-word: data mining, data clustering, database, density-based clustering, grid-based clustering, algorithm Density-based clustering algorithm can recognize arbitrary shapes. DBSCAN is a well-known algorithm that searches the neighboring objects of a region, and determines which objects belong to the same cluster [2]. However, DBSCAN requires a long search time. Therefore, IDBSCAN employs the sampling in the search scope to decrease the time cost [3]. FDBSCAN adopts quick merging to reduce redundant searching and raise efficiency [4].

Read the paper · More papers on PaperTik