A Deep Cross-modal Prompt Learning Network for Artificial Intelligence Generated Image Quality Assessment
Yang Lü, Shuangyao Han, Zilu Zhou, Zifan Yang, Gaowei Zhang, Shaohui Jin, Xiaoheng Jiang, Mingliang Xu · Displays · 2025
In recent years, multi-modal vision–language pre-trained models have been extensively adopted as foundational components for developing advanced Artificial Intelligence (AI) systems in computer vision applications. Previous approaches have advanced Artificial Intelligence Generated Image Quality Assessment (AGIQA) research via text-based or visual prompt learning, yet most methods remain constrained to a single modality (language or vision), overlooking the interplay between text and image. To address this issue, we propose a Deep Cross-Modal Prompt Learning Network (DCMPLN) for AGIQA. This model introduces a Multimodal Prompt Attention (MPA) module, employing multi-head attention to enhance the integration of textual and visual prompts. Furthermore, an Image Adapter module is incorporated into the visual pathway to extract novel features and fine-tune pre-trained ones using residual-style fusion. Experimental results on multiple generated image datasets demonstrate that the proposed method outperforms existing state-of-the-art image quality assessment models. • A Deep Cross-Modal Prompt Learning Network enhances AGIQA performance. • MPA module strengthens visual–text coupling via multi-head attention. • Image Adapter refines visual features with residual fusion and bottleneck.