A CNN-Transformer Hybrid Network for Multi-scale object detection
Jianhong Wu, Yingdong Ma · 2023
Recently, Transformer-based methods have been the main framework in various computer vision tasks. Vision transformers achieve object detection based on sequence of visual tokens, lacking the ability of extracting local context and dealing with scale variance. To tackle this problem, we propose a hybrid CNN-transformer model with multiple dual-branch transformer blocks in which transformer branch captures global dependences and the CNN branch enhances local context. As convolution branch and transformer branch pay attention to different-level information, we combine the output features of the CNN branch and transformer branch with adaptive weights calculated from the visual content. Moreover, instead of detecting objects from transformer outputs directly, we introduce a feature aggregation module to fuse different levels features and construct feature pyramid based upon these multi-level features. The proposed feature aggregation module alleviates semantic gap between high-level and low-level features. Experimental results on the MS COCO dataset show that our method significantly improves the performance of multi-scale object detection.