Selective Visual Relationship Detection with Adaptive Context Learning

Jian Tu, Feng Liu · 2024

Visual relation detection aims to describe the relationships between objects in a scene by using the form of a triplet. Existing methods not only suffer from the huge number of combinations of triples, but also make the context of most of the object pairs in a scene non-discriminative due to simply using the union box of the object pair. In this work, we propose a novel two-stage network. In the first stage, the multi-modal feature fusion strategy is used to complete the correlation detection of object pairs, which reduces the number of triples and enables the model to selectively focus on the relationship between strongly correlated objects in the scene. The object-directed context module (OCM) and global message passing (GMP) introduced in the second stage to adaptively learning the higher-order context of each object pair in the scenario. We evaluate our method on two widely used datasets: Visual Relationship Detection (VRD) and Visual Genome(VG) datasets. Experimental results verify the effectiveness of our method, e.g., the top 100 phrases detection recall improves from 40.01% to 41.21% on the VRD dataset.

Read the paper · More papers on PaperTik