MR-DSJ : Distance-Based Self-Join for Large-Scale Vector Data Analysis with MapReduce

Thomas Seidl, Sergej Fries, Brigitte Boden · 2013

Abstract: Data analytics gets faced with huge and tremendously increasing amounts of data for which MapReduce provides avery convenient and effective distributed programming model. Various algorithms already support massive data analysis on computer clusters but, in particular, distance-based similarity self-joins lack efficient solutions for large vector data sets though they are fundamental in many data mining tasks including clustering, near-duplicate detection or outlier analysis. Our noveldistance-based self-join algorithm for MapReduce, MR-DSJ, is based on grid partitioning and delivers correct, complete, and inherently duplicate-free results in asingle iteration. Additionally we propose several filter techniques which reduce the runtime and communication of the MR-DSJ algorithm. Analytical and experimental evaluations demonstrate the superiority over other join algorithms for MapReduce. 1

Read the paper · More papers on PaperTik