QHash: An efficient hashing algorithm for low-variance image deduplication

Xuan Li, Liqiong Chang, Xue Liu · 2021

Attributed to the widespread use of general artificial intelligence (AI), large-scale datasets have become a critical component for the success of AI-powered applications. While collecting a larger dataset is desirable in general, studies have revealed that many databases include duplicated images. The accumulation of redundant images will result in excessive resource usage and inefficient cloud storage utilization. To get rid of the duplicates, hashing-based methods have been developed for the problem of image deduplication. However, as demonstrated in our experiments, current approaches fail in dataset with small visual difference, such as medical images. To this end, we propose QHash, which achieves effective image deduplication on the low-variance dataset. QHash leverages Vector Quantized Variational AutoEncoder (VQ-VAE) to learn the data distribution in an unsupervised manner. In addition, hash sequences are implemented using integer-based tensor, enabling distance calculation and deduplication being processed in parallel on either CPU or GPU. Extensive experiments show that QHash outperforms other baseline approaches by at least 50%, while is about 23% memory efficient and 18% deduplication time speedup.

Read the paper · More papers on PaperTik