Adaptive multimodal deep prompt learning for blind image quality assessment

Yang Lu, Kunyan Lv, Zilu Zhou, Xiaoheng Jiang · Journal of Electronic Imaging · 2025

Blind image quality assessment aims to quantitatively evaluate the visual quality of images without requiring a reference to high-quality pristine images, making it a core task in computer vision. Recently, text-based prompt learning has shown great potential in adapting contrastive language-image pretraining (CLIP) models for evaluating the quality of natural images. Nonetheless, current unimodal approaches primarily focus on optimizing the language component of the CLIP model through prompt learning while overlooking the potential of prompt learning in the visual branch. To address this limitation, we propose adaptive multimodal deep prompt learning for image quality assessment (AdmpIQA), a innovative multimodal deep prompt learning framework. Specifically, AdmpIQA adapts the CLIP model by jointly optimizing learnable prompts in both the language and visual branches and introduces learnable prompts in each transformer of the encoder, enabling fine-grained adjustments to vision-language representations. In addition, we introduce a visual feature adaptation module designed to extract fine-grained features. This enhancement significantly improves the expressive capability of the visual branch, allowing it to more effectively capture the characteristics of image distortion. A wide range of experiments conducted across three IQA benchmark datasets confirms that AdmpIQA achieves superior performance compared with existing image quality assessment models.

Read the paper · More papers on PaperTik