Cross-Modal Attribute Insertions for Assessing the Robustness of Vision-and-Language Learning
Shivaen Ramshetty, Gaurav Verma, Srijan Kumar · 2023
The robustness of multimodal deep learning models to realistic changes in the input text is critical for their applicability to important tasks such as text-to-image retrieval and cross-modal entailment.To measure robustness, several existing approaches edit the text data, but do so without leveraging the cross-modal information present in multimodal data.Information from the visual modality, such as color, size, and shape, provide additional attributes that users can include in their inputs.Thus, we propose cross-modal attribute insertions as a realistic perturbation strategy for vision-and-language data that inserts visual attributes of the objects in the image into the corresponding text (e.g., "girl on a chair" → "little girl on a wooden chair").Our proposed approach for cross-modal attribute insertions is modular, controllable, and task-agnostic.We find that augmenting input text using cross-modal insertions causes state-of-the-art approaches for text-to-image retrieval and cross-modal entailment to perform poorly, resulting in relative drops of ∼ 15% in MRR and ∼ 20% in F 1 score, respectively.Crowd-sourced annotations demonstrate that cross-modal insertions lead to higher quality augmentations for multimodal data than augmentations using text-only data, and are equivalent in quality to original examples.We release the code to encourage robustness evaluations of deep vision-and-language models: https://github.com/claws-lab/ multimodal-robustness-xmai.