Improving Referring Expression Comprehension by Suppressing Expression-unrelated Proposals
Qun Wang, Feng Zhu, Xiang Li, Ruida Ye · 2023
The prevailing referring expression comprehension (REC) method is based on two stages: 1) detecting multiple candidate proposals using an object detector and 2) aligning proposals with text descriptions. Existing two-stage models have achieved significant improvements with excellent alignment modules. However, their alignment modules treat each proposal equally without filtering expression-unrelated proposals. These expression-unrelated proposals may interfere with the alignment modules’ ability to accurately reason about the target proposal. These proposals also increase the difficulty of training alignment modules. To solve this problem, we propose a new framework, that can suppress expression-unrelated proposals for alignment modules. Specifically, we first design a suppression module, which calculates the relevance scores between the proposals and the expression by using the semantics in the expression. Then, these scores are integrated into three basic alignment modules, including subject, location, and relationship modules, to guide each module to focus on the expression-related proposals. We conduct extensive experiments on three benchmark datasets, and the results show the effectiveness of our suppression module and performance improvement.