Type-Specific Modality Alignment for Multi-Modal Information Extraction

Shaowei Chen, Shuaipeng Liu, Jie Liu · IEEE Signal Processing Letters · 2024

Multi-modal information extraction aims to identify structured information, such as entities or relations between entities, from text with the help of visual clues. Although existing studies have achieved great progress, they mainly focused on modality interactions in the global space while neglecting fine-grained modality alignment under the semantic subspace specific to each entity type or relation type. To solve this problem, we propose a multi-space modality alignment method (MSMA) in this letter. The core of our model is a typespecific modality interaction module (TMI), which constructs a unique semantic subspace for each entity/relation type and independently performs type-specific modality alignments under each subspace. To enable mutual promotion between different types, a global modality integration module (GMI) is designed to learn the associations between different subspaces. Furthermore, we execute these two modules iteratively for high-level semantic fusion. Extensive experiments on three benchmark datasets show that our model significantly outperforms advanced methods.

Read the paper · More papers on PaperTik