Blind Image Quality Assessment With Multimodal Prompt Learning
Nhan Trong Luu, Chibuike Onuoha, Truong Cong Thang · 2023
Assessing the perception of visual content has long been a challenge in the field of computer vision. Numerous mathematical models have been created for image quality assessment (IQA). Despite the efficacy of these tools in quantifying visible distortions, their connection to human perception remains somewhat indirect. Particularly in capturing more abstract perception of visual content without any reference information (or blind image quality assessment), existing methods often rely on supervised learning with labeled data obtained through labor-intensive user studies. Recent research has deviated from traditional approaches by delving into the rich visual language embedded in Contrastive Language- Image Pre-training (CLIP) models. These studies aim to evaluate the quality perception of images without the need for explicit task-specific training. In a similar vein, this work introduces a new IQA model based on multimodal prompt learning (denoted MaPLe-IQA), for blind image quality assessment. Extensive experiments have been conducted on controlled IQA benchmarks. Findings demonstrate that the proposed quality model is capable of capturing perceptual assessments effectively without relying on any reference images. To our best knowledge, this is the first study that adopts the pretrained MaPLe backbone for the IQA tasks. The code for implementing this model is publicly accessible at https://github.com/luutn2002/mapleiqa.