Leveraging Deep Learning Technique with Meta-Data-Featured Images Captured by Drone for Real-Time Victim Positioning in Flood-Affected Areas
Trong Tuan Do, Dinh Duy Nguyen, Xuan Hai Ho, Van Duc Long Nguyen · 2025
Flooding has long been a major concern for nations frequently affected by natural disasters. The detection and localization of isolated victims for immediate temporary rescue have been a focal point for many researchers aiming to demonstrate the practical value of their solutions. However, accessing flood-affected areas poses significant challenges, as floods severely disrupt local transportation and damage infrastructure, rendering access nearly impossible. In such scenarios, the most feasible rescue approach is aerial intervention using Unmanned Aerial Vehicles (UAVs). Nevertheless, accurately detecting and localizing targets from a UAV's perspective is a significant challenge. UAVs operate at various altitudes and speeds, resulting in motion blur of targets against densely populated backgrounds. Additionally, target localization is made difficult due to the limited contextual understanding of the surroundings. To address these challenges, we propose training the YOLOv10 model to enhance the detection of small objects, integrated with a GPS-based target localization system. Furthermore, to enable the practical, real-world application of this solution, we introduce a communication infrastructure that facilitates the timely and effective deployment of both the detection and localization models. Field experiments have yielded promising results, with the artificial intelligence model achieving an Average Precision (AP) between 85–95%, and the localization model processing targets at an impressive speed of under 1 ms. The combined detection and localization approach demonstrated a positioning error of only 7–9 meters. Moreover, integrating the AI model and the localization algorithm into a low-latency communication system has shown promising results, providing near-instantaneous outputs with delays ranging from 0.6 to 0.9 seconds from the moment a target enters the frame.