Clusterers: a comparison of partitioning and density-based algorithms and a discussion of optimisations
David Breitkreutz, Kate Casey · ResearchOnline at James Cook University (James Cook University) · 2008
Though data mining is a relatively recent innovation, the improvements it offers over traditional data analysis have seen the field expand rapidly. Given the critical requirement for the efficient and accurate delivery of useful information in today's data-rich climate, significant research in the topic continues. Clustering is one of the fundamental techniques adopted by data mining tools across a range of applications. It provides several algorithms that can assess large data sets based on specific parameters and group related data points. This paper compares two widely used clustering algorithms, K-Medoids and Density-Based Spatial Clustering of Applications with Noise (DBSCAN), against other well-known techniques. The initial testing conducted on each technique utilises the standard implementation of each algorithm. Further experimental work proposes and tests potential improvements to these methods, and presents the UltraK-Medoids and UltraDBScan algorithms. Various key applications of clustering methods are detailed, and several areas of future work have been suggested.