Multi-modal Domain Adaptation for Text Visual Question Answering Tasks

Zhiyuan Li, Dongnan Liu, Weidong Cai · 2023

Domain adaptation aims to train a model on the labeled source data and unlabeled target data while improving the performance of the same model on the target domain. Recently, multi-modal domain adaptation is a relatively more challenging task compared with single-feature ones due to diverse forms of feature modalities. This paper aims to achieve transfer learning between two Visual Question Answering tasks: Text Visual Question Answering (Text-VQA) and Scene Text Visual Question Answering (ST-VQA). Compared to typical domain adaptation methods on single-modality analysis, the critical challenge for these tasks is that the input features contain complicated information from visual objects, question words, and optical character recognition (OCR) tokens in different forms. We consider a novel domain adaptation framework to facilitate knowledge transfer under the multiple modalities. Specially, the proposed Text-VQA domain adaptation framework (TDAF) consists of three important modules: (1) A mix-feature module fuses input features from multiple modalities and reduces the feature discrepancy for the cross-domain fused features. (2) A discrepancy discriminator focuses on keeping the source prediction correctly while detecting the ambiguous target samples by maximizing the discrepancy of probability outputs between two domains. (3) An self-supervised learning module is used for reducing information loss during representation learning by computing cosine similarity between collected inputs and decoded outputs. Through three joint modules, we successfully improve the performance of domain adaptation from Text-VQA to ST-VQA by raising 19.6% of accuracy and 18.2% of Average Normalized Levenshtein Similarity (ANLS). In addition, extra experiments were carried out to demonstrate the effectiveness of our proposed method compared with other SOTA approaches.

Read the paper · More papers on PaperTik