Multimodal-Enhanced Objectness Learner For Corner Case Detection In Autonomous Driving

Lixing Xiao, Ruixiao Shi, Xiaoyang Tang, Yi Zhou · 2024

Previous works on object detection have achieved high accuracy in closed-set scenarios, but their performance in open-world scenarios is not satisfactory. One of the challenging open-world problems is corner case detection in autonomous driving. Existing detectors struggle with these cases, relying heavily on visual appearance and exhibiting poor generalization ability. In this paper, we propose a solution by reducing the discrepancy between known and unknown classes and introduce a multimodal-enhanced objectness notion learner. Leveraging both vision-centric and vision-language multiple modalities, our semi-supervised learning framework imparts objectness knowledge to the student model, enabling class-aware detection. Our approach, Multimodal-Enhanced Objectness Learner (MENOL) for Corner Case Detection, significantly improves recall for novel classes with lower training costs. By achieving a $76.6 \% \mathrm{mAR}$-corner and $79.8 \%$ mAR-agnostic on the CODA-val dataset with just 5100 labeled training images, MENOL outperforms the baseline ORE by $71.3 \%$ and $60.6 \%$, respectively. The code will be available at https://github.com/tryhiseyyysum/ MENOL.

Read the paper · More papers on PaperTik