Few-Shot Multimodal Learning for Social Relation Understanding With Supervised Bi-Transformers
Hanqing Li, Niannian Chen · 2023
Multimodal social relationship understanding is a research field that integrates information from various sources, including visual and textual modalities, to identify and comprehend interpersonal relationships. Feature fusion is a critical factor in the task of extracting social relationships from multimodal sources, which significantly impacts recognition accuracy. However, the existing methods for multimodal social relationship understanding mainly rely on facial features in images and do not adequately consider the role of global image features in providing contextual information. Moreover, current methods lack in-depth exploration of Multimodal fusion. In contrast to traditional multimodal relationship extraction, the few-shot scenario task faces more semantic gap issues, such as inadequate cross-modal assistance and imbalanced relationships. To address these issues, this paper proposes a few-shot-based multimodal social relationship extraction method, namely SBT-MSRE. This method effectively combines global image features with facial features of head and tail entities and efficiently fuses textual and visual features to extract social relationships. We conducted extensive experiments on three challenging benchmark datasets, and the results demonstrate that our proposed method, SBT-MSRE, outperforms the current state-of-the-art methods in terms of accuracy.