VL-MFER: A Vision-Language Multimodal Pretrained Model With Multiway-Fuzzy-Experts Bidirectional Retention Network
Chen Guo, Xinran Li, Jiaman Ma, Yimeng Li, Yuefan Liu, Haiying Qi, Li Zhang, Yuhan Jin · IEEE Transactions on Fuzzy Systems · 2024
Vision-language pretrained models have achieved significant success across various tasks. However, the lack of interpretability limits the application of multimodal models, especially for those that require the security of systems, data, and users. A key challenge in enhancing the interpretability of the models is the tradeoff between transparency and model performance in terms of accuracy and computing efficiency. We propose a vision-language multimodal pretrained model called multiway-fuzzy-experts bidirectional retention network (VL-MFER), which is designed to effectively interpret model decision processes while enhancing the consistent mapping between text and image features. We first propose a bidirectional retention network to handle and integrate the cross-modal high-dimensional data, which can effectively enhance both the performance and inference efficiency. Then, to improve the interpretability of vision-language pretrained models such as the unified vision-language pretraining with mixture-of-modality-experts, we propose a multiway fuzzy experts pool based on multiple different layered deep neuro-fuzzy systems for diverse downstream tasks. After training on diverse datasets consisting of images, text, and paired image-text data, we fine-tuned models for multiple vision and vision-language tasks. Experimental results show that VL-MFER performs well in all performance metrics of downstream tasks, and VL-MFER leads to enhancements in transparency compared to other methods, while improving computational efficiency and reducing inference time by up to 13%.