A Vision-Language Pre-training model based on Cross Attention for Multimodal Aspect-based Sentiment Analysis

HengRui Hu · 2024

Multimodal aspect-based sentiment analysis (MABSA) focuses on extracting unique aspect term within multimodal informations and subsequently analysising sentiment associated with this aspect. Existing studies have typically used either a separate pipeline method or a uniform transformer. However, these approaches are inadequate to clearly and efficiently integrate the alignment between distinct modes. To resolve these constraints, we suggest a Vision-Language Pre-training model based on cross attention for MABSA named VLPCA, which applies a novel Multi-head cross attention to capture textual and visual features for better representation of visual-language interactions. Furthermore, we design two subtasks and propose a novel unsupervised joint training approach based on contrastive learning to enhance the effectiveness of the proposed model. Extensive experiments and ablation studies show that our model consistently performs better than the current methods.

Read the paper · More papers on PaperTik