Deep learning for UAV thermal bird detection: Benchmarking YOLOv8–YOLOv26
Anastasiia N. Safonova, Arnau Campanera, João Paulo Silva, Diego Felipe Navarro Sarró, Marcelino Cardalliaguet Guerra, Joan S. Serra-Sagrista, Andrea Santangeli · Remote Sensing Applications Society and Environment · 2026
Unmanned Aerial Vehicles (UAVs) equipped with thermal imaging sensors offer a valuable solution for automated wildlife monitoring, particularly in challenging environments where visibility is limited. However, accurately detecting small and thermally subtle targets like birds remains difficult due to environmental complexity, low contrast, and variation in object size within images. YOLO (You Only Look Once) is one of the most widely adopted Deep Learning (DL) frameworks for object detection. Its applications range from traffic monitoring to medical imaging. Recently, YOLO has also been used for ecological research, including bird detection. However, its effectiveness for wildlife monitoring in real-world field conditions remains largely untested. This study benchmarks 31 YOLO models (v8–v26, nano to x -large variants) for detecting little bustards ( Tetrax tetrax) in independently constructed UAV-captured thermal images using various data augmentation techniques to enhance model robustness. A curated dataset of 2,942 annotated images spanning multiple seasons and landscapes in southwestern Spain was used for training and evaluation. Results show a clear stratification in performance across model sizes, with larger architectures generally achieving higher detection accuracy and more stable convergence. The highest validation mAP@50–95 was achieved by YOLOv8x (0.693), followed by YOLOv10x (0.684) and YOLOv9c (0.674), with the new YOLOv26n also performing strongly (0.669). On the test set, YOLOv11n achieved the best performance (0.631), closely followed by YOLOv12l (0.630) and YOLOv26m (0.627). Despite achieving consistently high precision and recall (≈0.96–0.98) across all models, clear differences in generalization emerged. Top-performing architectures on the validation set such as YOLOv8x, YOLOv10x and YOLOv9c, as well as the recently introduced YOLOv26n, did not necessarily maintain their ranking on the test set. For instance, YOLOv11n and YOLOv12l achieved the highest test mAP@50–95 scores, while YOLOv26m also performed strongly despite not leading on the validation set. Although smaller variants are computationally efficient, they exhibited lower mAP@50–95 values and were more prone to false detections. These findings reinforce the trade-offs between performance and efficiency across the YOLO family, and highlight the importance of selecting architectures that balance accuracy, robustness and computational cost for reliable UAV-based thermal wildlife monitoring.