A Vision Enhanced Framework for Indonesian Multimodal Abstractive Text-Image Summarization
Yutao Song, Nankai Lin, Lingbao Li, Shengyi Jiang · 2024
Multimodal abstractive summarization (MAS) is a technique that generates a brief summary by processing input text and images. While preceding investigations on MAS have prioritized the utilization of visual features to amplify the quality of summaries, such advancements have predominantly been realized within high-resource languages, most notably English and Chinese. However, in the case of resource-scarce languages like Indonesian, the research and available resources pertaining to multimodal abstractive summarization remains constrained. In addition, the heterogeneity between visual and textual features may impact the quality of summary generation. Therefore, it is crucial to investigate vision-enhanced generative models to improve summary quality. To address the problem of insufficient resources of Indonesian MAS, we constructed the E-Liputan dataset, which is a summary-guided multimodal generative summarization dataset in Indonesian. We employed a two-stage methodology: first, we utilized a mask strategy to effectively address text denoising, thereby facilitating the pre-training of the visual encoder. Second, we fine-tuned an end-to-end multimodal summarization model and proposed a summary-guided multimodal interactive co-attention learning fusion, which facilitated the seamless integration and fusion of modalities within the model. Through the employment of these methodologies, the proposed multimodal model effectively acquired a richer repertoire of visual feature information oriented towards summarization, ultimately leading to enhanced precision and accuracy in generating summaries. We conducted extensive experiments on the E-Liputan dataset and found that our model outperformed the baseline models. Our findings suggest that investigating vision-enhanced generative models for MAS can significantly improve summary quality, particularly in resource-scarce languages.