Enhancing Intermodal Interaction for Unified Vision-Language Understanding and Generation
Qin Yang, Huiming Xie, Yujie Li, Benying Tan, Shuxue Ding · Data Intelligence · 2025
The majority of vision-language pre-training (VLP) models rely on pre-trained object detectors, which incur high costs and restrict the recognition of object classes. Additionally, their encoder-based structures hinder their ability to perform text generation tasks effectively. To mitigate these challenges, we propose a Detector-free Vision-and-Language Pre-training (D-VLP) model designed to bolster intermodal interaction for unified understanding and generation tasks. Our D-VLP model employs a co-modality decoder equipped with a fused multi-attention self-attention module, enhancing feature fusion and information alignment between images and text. It is pre-trained using a novel Prefix Masked Language Modeling (prefixMLM) approach, leveraging the strengths of masked language modeling and unidirectional language modeling, which enables bidirectional processing and autoregressive token generation. Extensive experiments demonstrate that D-VLP surpasses state-of-the-art models in vision-language tasks, highlighting its superior performance and adaptability across various image-text tasks with minimal adjustments.