MODCL: multi-modal object detection with end-to-end contrastive learning in indoor scene

Zixu Lan, Fang Deng, Angang Zhang, Zhongjian Chen · 2024

In recent years, research on multi-modal object detection has garnered significant attention due to the comprehensive information obtained from multi-modal data. However, most object detection studies focus on outdoor scenarios, such as autonomous driving, with relatively few investigations addressing the characteristics of indoor scenes. In this paper, we identify the shortcomings of similar object detection performance in indoor scenes and propose MODCL: Multi-modal Object Detection with End to End Contrastive Learning in Indoor Scene. Within MODCL, we focus on two aspects of improvement: first, the fusion of multi-modal context based on the mapping relationship between point clouds and images; second, the incorporation of supervised contrastive learning in an end-to-end manner, eliminating the need for pre-training. Furthermore, we conducted experiments on the SUN RGB-D dataset, and the results indicate that MODCL outperforms existing detection methods that utilize both point clouds and images compared to those that rely solely on point clouds.

Read the paper · More papers on PaperTik