AMAF-YOLO: dynamic cross-region attention and multi-scale fusion for small object detection

Qiang Sun, Jinhua Zhang, Shurong Zeng, Jun Hu, Kangmei Li · Nondestructive Testing And Evaluation · 2025

Small object detection in complex backgrounds is crucial in computer vision. However, existing approaches are often limited by low resolution, high object density, and background clutter, leading to persistently high false negative and false positive rates. To address this, we propose AMAF-YOLO, a lightweight detector based on YOLOv12, incorporating a dynamic cross-spatial attention mechanism and multi-scale feature fusion. The model introduces a Cross-Stage Network with Lightweight Global Context (CSP_LGC), which maintains efficient local feature extraction while capturing cross-region contextual relationships, significantly reducing complexity. Additionally, a Multi-scale Augmented Feature structure (AMAF) combines cross-layer serial residual modules with The dual cross-fusion attention mechanism (DCFA) to formulate a dynamic cross-region attention structure, which effectively aggregates multi-scale features and improves the granularity feature representations across different hierarchical levels. An Enhanced Residual Gate (ERG) module uses a dual-branch design to jointly capture fine details and wide-context features, boosting small object detection. On VisDrone2019, AMAF-YOLO improves accuracy by 6.0%, recall by 4.3%, mAP0.5 by 5.8%, and mAP0.5–0.95 by 4.1%, while reducing parameters count by 54.7% and model size by 50.9% compared to YOLOv12n. Further validation on AI-TOD, SRSDD-V1.0, and BSTLD datasets confirmsthe method’s strong generalisation capability and performance superiority.

Read the paper · More papers on PaperTik