Generalized DINO: DINO via Multimodal Models for Generalized Object Detection

Kashing Yuen, Jianpeng Zou, Kaoru Uchida · 2024

Referring Expression Comprehension (REC) is a task in the realm of vision and language, aiming to identify objects in images based on provided descriptions. Classic REC methods, however, face challenges in handling expressions involving multiple targets or empty scenarios. In this paper, we study the limitations of existing REC methods, particularly in the context of Generalized Referring Expression Segmentation (GRES). In response, we propose Generalized DINO, a model that extends Transformer-based detectors by incorporating Region-Image Cross Attention (RIA) and Region-Language Cross Attention (RLA) mechanisms. This approach enables the detector to support arbitrary numbers of target object detection, overcoming the constraints of traditional REC methods. Comprehensive experiments on widely-used datasets such as RefCOCO/+/g and the GRES benchmark gRefCOCO showcase the superior performance of Generalized DINO in GRES tasks. The model outperforms even the robust RELA model, demonstrating a significant stride in handling expressions with multiple targets or empty scenarios. Our findings underscore the efficacy of Generalized DINO in enhancing the robustness and flexibility of REC models, contributing to multimodal information processing. The model's ability to handle complex language expressions involving multiple objects positions it as a valuable asset in applications like human-computer interaction and visual question answering.

Read the paper · More papers on PaperTik