Unified Multi-Modal Multi-Task Joint Learning for Language-Vision Relation Inference

Wenjie Lu, Dong Zhang · 2022 IEEE International Conference on Multimedia and Expo (ICME) · 2022

The relationships between language and vision are valuable for natural language processing and computer vision research, where the text and image data are employed to develop computing techniques for image caption or visual grounding. Although the existing studies have been engaged in language- vision relation inference (LVRI), they are limited to low- resource of the task itself. In this paper, we mainly focus on LVRI upon text-image pairs in Twitter with a unified multi- modal multi-task joint learning approach. Different from the conventional multi-modal multi-task learning approach within the same multi-modal dataset, we leverage a relevant multi - modal task on the external dataset as an auxiliary task to facilitate the LVRI task. Systematic experiments demonstrate the effectiveness of our proposed multi-modal multi- taskjoint learning approach.

Read the paper · More papers on PaperTik