PhotoCluster - A Multi-clustering Technique for Near-duplicate Detection in Personal Photo Collections

Vassilios Vonikakis, Amornched Jinda-apiraksa, STEFAN S. WINKLER · 2014

This paper presents PhotoCluster, a new technique for identifying non-identical near-duplicate images in personal photo collections. Contrary to existing methods, PhotoCluster estimates the probability that a pair of images may be considered near-duplicate. Its main thrust is a multiple clustering step that produces a non-binary near-duplicate probability for each image pair, which exhibits correlation with the average observer opinion. First, PhotoCluster partitions the photolibrary into groups of semantically similar photos, using global features. Then, the multiple clustering step is applied within the images of these groups, using a combination of global and local features. Computationally expensive comparisons between local features are taking place only on a limited part of the library, resulting in a low overall computational cost. Evaluation with two publicly available datasets show that PhotoCluster outperforms existing methods, especially in identifying ambiguous near-duplicate cases.

Read the paper · More papers on PaperTik