The application of K-means clustering algorithm based on Hadoop

Yurong Zhong, Dan Liu · 2016

Spatial data is different from the general data, it not only contains some kind of property information of space feature, but also has the spatial feature of space or location. The spatial clustering analysis can be divided into two broad categories. Category from GIS theory and technology tools, according to an object of spatial geographical coordinates, cluster as an object of the spatial proximity base rather than as a clustering property similarity. Category from the application of GIS and geoscientific research angle, the general is feature of space object's properties by using traditional clustering methods and clustering analysis, but ignores the geographical coordinate information of the space, and would not considers the spatial proximity of objects. In order to solve this kind of geographical position and property feature of double meaning, spatial data mining based on K - means spatial clustering algorithm combining geographic location and property feature, practise unified entity properties of spatial proximity and similarity. Considering the sharping increasing scale of spatial data, the realization of the K - Means spatial clustering based on graphs parallel algorithm.

Read the paper · More papers on PaperTik