A Fast Density-Grid Based Clustering Method
Daniel Brown, Arialdis Japa, Yong Shi · 2019
Clustering algorithms are a large area of focus in the world of data analytics, especially in the current age of big data sets that are large, multidimensional, and grow quickly. We also live in an age where parallel and distributed computing are at the forefront of any form of data analysis. Because of these two ideas, this paper discusses an algorithm that was designed to scale well with big data sets for fast run time, while also being highly scalable for a potential parallel version. This algorithm, the Fast Density-Grid Clustering Algorithm, works by dividing the data space into a grid structure and then assigning a density measurement to each grid cell. These spaces are then merged based on their densest neighbor in order to cluster the entire space. The experimental results for the serial version demonstrate that the algorithm is generally comparable to DBSCAN in accuracy, while also having a lower run time.