Explainable AI for Image Aesthetic Evaluation Using Vision-Language Models

Supatta Viriyavisuthisakul, Shun Yoshida, Kaede Shiohara, Ling Xiao, Toshihiko Yamasaki · 2025

Evaluating image aesthetics is inherently subjective and has traditionally relied on the expertise of human evaluators. Vision-Language models, such as Contrastive Language-Image Pre-Training (CLIP), present a new paradigm for assessing the visual features and descriptions of images, thus enabling more interpretable aesthetic evaluations. Recently, CLIP-based image quality assessment (CLIP-IQA) has emerged as a method to quantify quality and abstract perception in images using an antonym prompt pairing strategy. Despite achieving a high correlation with human aesthetic judgment, questions persist regarding the relevance of features that align with human perception. In this study, we investigate the significance of image features derived from various paired prompts. Each prompt pair is encoded into feature vectors using a text encoder, while images are encoded in a similar manner using an image encoder. To predict quality scores, we use Light Gradient Boosting Machine (LightGBM) as a regressor. After training, SHapley Additive exPlanations (SHAP) values are computed for each feature, enabling us to evaluate the contribution of individual prompt elements. In this study, a multimodal large language model (MLLM) is applied to generate the linguistic explanations of images. Our results yield Spearman’s rank correlation coefficient (SROCC) and Pearson linear correlation coefficient (PLCC) scores of 0.762 and 0.785, respectively. Furthermore, we explore advanced prompting strategies, uncovering deeper insights into the IQA scoring mechanism.

Read the paper · More papers on PaperTik