An efficient similarity join approach on large‐scale high‐dimensional data using random projection

Youzhong Ma, Ruiling Zhang, Shijie Jia, Yongxin Zhang, Xiaofeng Meng · Concurrency and Computation Practice and Experience · 2019

Summary Similarity join on large‐scale high‐dimensional data faces major challenges because of the data scale and the cure of dimensionality. Random projection with p‐stable distribution can reduce the high‐dimensional data form d‐dimension to k‐dimension (k ≪ d), the distance of the data in k‐dimensional space can be used to filter out as many data pairs as possible at relative low cost. Based on the above idea, we proposed two novel approaches to deal with large‐scale high‐dimensional data similarity join: projection‐based similarity join (PromSimJ) algorithm and projection space partitioning–based similarity join (ProSPSimJ) algorithm. The comprehensive experiments were performed to test the performance of the above methods. We also compared the performance of the above methods with that of the naive method block nested loop join. The final experimental results prove that our approaches have much better performance and good scalability.

Read the paper · More papers on PaperTik