Improvement of k-nearest neighbor algorithm based on double filtering

Chun Jie, Zheng Sheng Ding · 2020 5th International Conference on Mechanical, Control and Computer Engineering (ICMCCE) · 2020

Aiming at the shortcoming that the efficiency of k-nearest neighbor algorithm (kNN) greatly decreases with the increase of the number of samples, a two-screen k-nearest neighbor algorithm is proposed. Its purpose is to increase the calculation speed of the algorithm while ensuring that the accuracy rate does not decrease, so that the improved algorithm can handle big data problems. Firstly, large-scale spectral clustering based on landmark representation (LSC algorithm) is used to cluster the data to find the cluster closest to the point to be measured; secondly, triangle inequality is used to process the data set close to the point to be measured and far from the nearest cluster center. The data after these two screenings is then used as the final training sample of the k-nearest neighbor algorithm. It can improve the diversity and effectiveness of data on the basis of reducing the number of training samples. Finally, it was verified on a public data set and compared with the traditional k nearest neighbor algorithm. The results show that the speed of the algorithm has been improved, the calculation amount has been reduced by 49%, and the accuracy has been maintained at about 95%.

Read the paper · More papers on PaperTik