Faster Transformer-DS: Multiscale Vehicle Detection of Remote-Sensing Images Based on Transformer and Distance-Scale Loss

Jiahuan Zhang, Hengzhen Liu, Yi Zhang, Menghan Li, Zongqian Zhan · IEEE Journal of Selected Topics in Applied Earth Observations and Remote Sensing · 2023

Vehicle Detection (VD) on Remote-Sensing (RS) images has gained impressive achievements these years, mainly thanks to the development of popular learning-based object detection architectures (e.g., Faster R-CNN, YOLO series, etc.). However, for RS images, muti-scale vehicle detection with tiny size still remains challenging. Particularly, vehicles as tiny objects typically contain only a few pixels with very rare information for model training and validation, which can result in inaccurate localization and intricate classification. In this paper, we present a new detection model called Faster-Transformer-DS, in which two improvements are proposed and discussed: (1) instead of CNN based ResNet50, a Transformer-based backbone—PVTv2-b0 that can extract global context information, is investigated as feature extraction backbone for both object classification and bounding box regression. (2) In contrast to the conventional paradigm-based and IoU (Intersection over Union)-based loss functions, to further fine-tune the localization of predicted bounding box, we proposed a novel distance-scale loss (DS Loss) function—the distance loss is directly related to the absolute value of the length and width of the bounding box together with its center position, while the scale loss refers to the geometric shape ratio between the bounding box and the ground truth box. To demonstrate the efficacy of our model, comprehensive experiments are conducted to show that both the proposed methods can boost the performance of multi-scale vehicle detection on RS images and the best result, all of which are confirmed by experimental data indicators and visual detection results. Additionally, several state-of-the-art methods are compared and consequently, our model is notably superior to many common detection models.

Read the paper · More papers on PaperTik