Unified Image Similarity Detection Using Neural Networks and Feature Metrics

Ch. Prathima, G. Mahesh Babu, A. Koushik, P. Arjun, S. Manoharra, M. Mahesh Babu · 2025

With the increase in the amount of image information in every field, the ability to find and eliminate duplicate or near-duplicate images among the huge number of copies stored in large databases becomes an even more vital task. This project targets determining and removing duplicate/nearly duplicate copies and keeping original images without any alterations only. Although YOLO, VGG16, InceptionV3, Alex Net, Reset, and Zente CNNs indeed did well in performing image classification and detecting image similarities, they still pose some limitations in the area of distortion and noise that are introduced through image-processing attacks. These limitations are very much evident in practical scenarios when images suffer from general attacks like blurring, noise addition, scaling, rotation, and compression, which make the accurate discrimination of original and distorted images quite difficult for the CNN models. In such a scenario, the feature matrix integrated with multiple measures of similarities, for example, SSIM, Histogram Similarity, Hash Similarity, and Key point Similarity, has been used in this project. It helps in comparing the images in a much-accurate manner even with distortions or noisy ones. The proposed system of this project employs an iterative process which first applies a series of image transformations or attacks such as Median Filter, Wiener Filter, Gaussian Filter, JPEG Compression, Motion Blur, and so on for simulations of real-world image variations. These include some of the usual distortions that may arise in the course of the acquisition, processing, and transmission of images. By using such transformed images, the synergistic use in the proposed system of the CNN model and the feature matrix in detecting duplicate and near-duplicate images becomes possible. One of the most prominent new features of the proposed system includes eliminating all versions of near-duplicated and processed images.

Read the paper · More papers on PaperTik