RETR: End-To-End Referring Expression Comprehension with Transformers

Rui Yang · 2022 19th International Computer Conference on Wavelet Active Media Technology and Information Processing (ICCWAMTIP) · 2022

Referring Expression Comprehension (REC) is a basic and challenging task to identify the referred region given a language expression. However, existing two-stage or one-stage methods suffer from the region proposals, the limited range of visual context and the incomplete cross-modal alignment. To address these problems, we propose a simple yet effective one-stage model, termed REC TRansformer (RETR), which is trained end-to-end. Different from the manually designed multi-modal fusion, RETR adopts a transformer decoder with alternately stacked self-attention and cross-attention layers to capture the global visual context and establish the detailed visual-linguistic correspondence. Moreover, we utilize multiple learnable tokens to obtain diverse yet complementary region representations to give the accurate prediction. Extensive experiments are conducted on four datasets and RETR achieves the state-of-the-art performance.

Read the paper · More papers on PaperTik