Optimal hash functions for approximate closest pairs on the n-cube

Daniel M. Gordon, Victor S. Miller, Peter Ostapenko · arXiv (Cornell University) · 2008

One way to find closest pairs in large datasets is to use hash functions. In recent years locality-sensitive hash functions for various metrics have been given: projecting an n-cube onto k bits is simple hash function that performs well. In this paper we investigate alternatives to projection. For various parameters hash functions given by complete decoding algorithms for codes work better, and asymptotically random codes perform better than projection.

Read the paper · More papers on PaperTik