MASNet: a novel deep learning approach for enhanced detection of small targets in complex scenarios
Zhenwen Zhang, Ya‐Yun Yang, Xianzhong Jian · Measurement Science and Technology · 2025
Abstract Accurate object detection plays a critical role in image processing, especially for identifying small targets in complex backgrounds. However, factors such as degraded image quality and extended viewing distances impede detection performance, compromising both accuracy and multi-target recognition. To address these challenges, we propose the Multi-Attention Spatial Network (MASNet), a streamlined object detection model that leverages attention mechanisms to enhance target features while suppressing noise interference in cluttered environments. MASNet introduces Space-to-Depth Convolution (SPDConv) to enrich the feature map depth while preserving spatial resolution, thereby improving local detail extraction and reducing computational overhead. Additionally, Global-Local Spatial Attention (GLSA) is integrated into the backbone to emphasize critical features across both global and local scales, enhancing fine-grained feature representation. Furthermore, Sparse Multi-Scale Attention (SMSA) is incorporated into the detection head to refine multi-scale feature learning, while the Efficient IoU (EIoU) loss function is employed to accelerate convergence, improve regression accuracy, and boost overall detection performance. Experimental results demonstrate that MASNet achieves a mAP of 52.2 % on the VisDrone2019 test set and 62.8 % on the validation set, significantly enhancing detection accuracy for aerial images. Moreover, MASNet attains a mAP of 89.7 % on the AFO dataset, 42.8 % on the UAVDT dataset, and 85.1 % on the UAVOD-10 dataset. We confirmed MASNet’s superior performance through comprehensive comparative evaluations across multiple datasets, demonstrating its robustness and efficiency in small target detection under diverse and challenging imaging conditions.