Object detection using deep learning (Object detection using deep learning on unmanned aerial vehicles)

Hongkai Lin · DR-NTU (Nanyang Technological University) · 2026

The growth of Object Detection of Unmanned Aerial Vehicles (UAVs) has significantly expanded the applications of aerial computer vision. However, it brings out the challenges as targets like cars and pedestrians appear extremely small, blurry and crowded. This Final Year Project investigates these challenges by evaluating and adapting two advanced artificial intelligence models on the VisDrone dataset: a lightweight single-stage detector (YOLO11n) and multi-model Vision-Language Model (Grounding DINO). The project was carried out in two main steps. The first step was preparing the VisDrone dataset and converting its labels into formats that each model could process. The second step was fine-tuning the pre-trained models using these newly formatted datasets. The YOLO11n model was trained to prioritize speed and basic accuracy for small objects. Meanwhile, Grounding DINO was being fine-tuned to learn how to perfectly match the complex drone images with our text descriptions. The final benchmark results show a clear trade-off between speed and accuracy. The fine-tuned Grounding DINO model proved to be highly accurate, achieving an mAP@50 score of 0.536. It is excellent at finding tiny, distant objects but runs slower, processing at about 9 FPS. On the other hand, the fine-tuned YOLO11n model had a lower accuracy score (mAP@50 of 0.344) but was incredibly fast, running at 60 FPS on an windows personal PC. Based on these results, this project concludes that drones require a two-part deployment strategy. YOLO11n is the ideal choice to be installed directly on the drone for fast, real-time flying and obstacle avoidance. Grounding DINO is the essential tool for ground computers to analyze the recorded video after the flight, providing maximum accuracy for tasks like search-and-rescue mapping and traffic counting.

Read the paper · More papers on PaperTik