Duplicate Image Detection in Large Scale Databases
Pratim Ghosh, Elisa Drelie Gelasca, K.R. Ramakrishnan, Bangalore S Manjunath · Statistical science and interdisciplinary research · 2008
We propose an image duplicate detection method for identifying modified copies of the same image in a very large database. Modifications that we consider include rotation, scaling and cropping. A compact 12 dimensional descriptor based on Fourier Mellin Transform is introduced. The compactness of this descriptor allows efficient indexing over the entire database. Results are presented on a 10 million image database that demonstrates the effectiveness and the efficiency of this descriptor. In addition, we also propose extension to arbitrary shape representations and similar scene detection and preliminary results are also included. Automated robust methods for duplicate detection of images/videos is getting more attention recently due to the exponential growth of multimedia content on the web. The large quantity of multimedia data makes it infeasible to monitor them manually. In addition, copyright violations and data piracy are significant issues in