CMMS-DETR: Cross-Modal Multiscale Feature Fusion Network for UAV Remote Sensing Detection

Yin Wang, Jiliang Mu, Qihong Yang, Guanglei Gao, Jian He, Junbin Yu, Xiujian Chou · IEEE Journal of Selected Topics in Applied Earth Observations and Remote Sensing · 2026

Target detection based on visible-infrared images is a key technology for remote sensing detection and tracking using unmanned aerial vehicles (UAVs). However, this technology still faces many challenges due to the inherent characteristics of measured targets and the captured images. For example, the scale variation of the measured target in the image is too large, the expression of target details and textures is insufficient, and the feature differences of different modal data seriously affect the accuracy of target detection. In this article, a cross-modal multiscale feature fusion network architecture based on RT-DETR (CMMS-DETR) for UAV target detection is proposed to solve these challenges. In order to improve the ability of multiscale target feature extraction, the proposed multiscale residual feature extraction module adds ParNet to ResNet to increase the network receptive field. In order to extract the multidomain detailed features of the measured target and deepen the multiscale features, a multiscale mixed channel attention module is designed, which extracts the multiscale and multichannel global features through two branches in parallel, and strengthens the selection of key features. Aiming at the problem of redundancy and insufficient fusion of multimodal features, a cross-modal feature fusion module is proposed, which uses channel-spatial dual attention to generate dynamic weights to complement detailed features and suppress irrelevant noise. The mean average precision of the proposed CMMS-DETR on DroneVehicle, LLVIP and antiUAV datasets are 81.40%, 97.30%, and 98.80%, respectively. A large number of experiments verify the advancement of this architecture.

Read the paper · More papers on PaperTik