Vision-Language Collaboration Enhancement for Unbiased Scene Graph Generation

Chenshu Wang, Jiying Wu, Cong Du, Gaoyun An · 2024

Scene Graph Generation (SGG) aims to detail the visual contents of images and present them as a compact summary graph. Although several approaches have made great progress in SGG, it’s still faced with the long-tailed distribution of predicates. Existing debiasing methods that focus on tail classes achieve a lot but lead to the performance degradation of head ones to some degrees. In this paper, we consider imperfect predicate feature representation is a key factor affecting the detection results of both head and tail predicates. Therefore, we focus on generating better predicate features targeting at enhancing detection capabilities for both head and tail ones. We propose a vision-language collaboration enhancement framework to obtain enhanced predicate features by fusing multiple features and context. Within this framework, we further propose an adaptive residual feature enhancement module to ensure the completeness of multi-scale features. Extensive experiments on VG dataset confirm the effectiveness of our method in improving overall performance as well as the predicate detection performance of both head and tail classes.

Read the paper · More papers on PaperTik