Multi-scale Fusion with Context-aware Network for Object Detection
Hanyuan Wang, Jie Xu, Linke Li, Ye Tian, Du Xu, Shizhong Xu · 2018
Almost all of the state-of-the-art object detectors employ convolutional neural network (CNN) to extract feature. However, how to fully utilize spatial information is a challenge. In this paper, we propose an effective framework for object detection. Our motivation is that multi-scale representation and context are extremely important for object detection. For multi-scale representation, our mothed combines hierarchical feature maps to a fusion map, which has abundant spatial information and high-level semantics. For context, we exploit spatial information by stacking multi-region feature maps. The network is learned end-to-end, by minimize an objective function. Our network achieves competitive results, 75.9% mAP on PASCAL VOC 2007, 72.0% mAP on PASCAL VOC 2012 and 23.2% mAP on MS COCO. The speed of the network is 10 fps. Our studies demonstrate that multiscale representation and context can further improve performance of object detection.