Improved shared nearest neighbor clustering algorithm

Shengyi Jiang · Computer Engineering and Applications Journal · 2011

Clustering is a method of unsupervised learning in machine learning,the typical task of which is to discovery naturalclusters present in the data.The shared nearest neighbor algorithm is one of the most efficient clustering algorithm which can handle datasets of different sizes,shapes and densities.But there are still some shortages about the algorithm.SNN can’t handle large dataset because of its high complexity.There are no definite methods about threshold of the algorithm. SNN can not process databases with mixture attributes.This paper improves the SNN algorithm to process the data with categorical attributes,gives a simple definite method to select threshold of the algorithm.The time complexity of the improved algorithm is nearly linear with the size of dataset and can be used to large dataset.The experimental results on real datasets and synthetic datasets show that the improved algorithm is effective and practicable.

Read the paper · More papers on PaperTik