Research on K nearest neighbor join for big data

Ji Jiaqi, Yeong-Jee Chung · 2017

K Nearest Neighbor Join (KNN Join) is a primitive operation widely adopted by many data mining applications. As a combination of the k nearest neighbor query and the join operation, KNN Join is a computationally intensive algorithm; however, with the increase of data volume and data dimension, the results can't be obtained within acceptable time when this algorithm runs on a single machine. Consequently, on the basis of Spark, a new approach that employs Locality-Sensitive Hashing (LSH) is proposed. The LSH algorithm first maps similar objects onto the same bucket, which can reduce the set of k nearest neighbors; then the distance of objects in the cluster, can be calculated based on Spark. The experimental results show that this proposed approach is accurate and effective for high dimensional big data.

Read the paper · More papers on PaperTik