Alignment of Perceptual Similarity Metrics with Human Perception
Abhijay Ghildyal · 2025
Perceptual similarity metrics are used for quantitatively evaluating the similarity between two images as it would appear to human perception. These metrics aim to mimic the human visual system, providing a more accurate assessment of visual similarity. Such visual assessments are considered to be more advanced than simple pixel-wise comparisons such as ℓp norm distances. Thus, a human-like assessment of visual similarity, makes the metrics valuable for applications in image compression, restoration, and enhancement, where evaluating perceptual quality is crucial. Perceptual similarity metrics have progressively become more correlated with human judgments on perceptual similarity; however, despite recent advances, the addition of an imperceptible distortion can still compromise these metrics. This dissertation investigates how the magnitude of specific perturbations applied to visual stimuli affects the responses of low-level perceptual similarity metrics, with the goal of determining how imperceptible distortions impact the reliability of these metrics. We begin by investigating the robustness of perceptual similarity metrics against an imperceptible geometrical distortion that can occasionally occur during image acquisition and preprocessing. Existing perceptual similarity metrics assume an image and its reference are well aligned. As a result, these metrics are often sensitive to a small alignment error that is imperceptible to the human eyes. In this dissertation, we first study the effect of small misalignment, specifically a small shift between the input and reference image, on existing metrics, and accordingly develops a shift-tolerant similarity metric. We build upon LPIPS, a widely used learned perceptual similarity metric, and explores architectural design considerations to make it robust against imperceptible misalignment. Specifically, we study a wide spectrum of neural network elements, such as anti-aliasing filtering, pooling, striding, padding, and skip connection, and discuss their roles in making a robust metric. Based on our studies, we develop a new deep neural network-based perceptual similarity metric. Our experiments show that our metric is tolerant to imperceptible shifts while being consistent with the human similarity judgment. We further extend our investigation by systematically evaluating the robustness of these metrics to imperceptible adversarial perturbations. We call these perturbations adversarial, as they are deliberately crafted by an attacker with malicious intent. Following the two-alternative forced-choice experimental design with two distorted images and one reference image, we perturb the distorted image closer to the reference via an adversarial attack until the metric flips its judgment. We first show that all metrics in our study are susceptible to perturbations generated via common adversarial