Ensemble learning and EigenCAM-based feature analysis for improving the performance and explainability of object detection in drone imagery
Gargi Joshi, Amey Joshi, Mranmay Shetty, Rahee A. Walambe, Ketan V. Kotecha, Fabio Scotti, Vincenzo Piuri · Discover Applied Sciences · 2025
Object Detection (OD) is an essential task in computer vision, which involves identifying and localizing objects in images or videos. The ability to detect objects automatically is crucial in surveillance, agriculture, infrastructure inspection, and search and rescue operations [ 1 ]. Over the years, deep learning-based techniques [ 2 ], particularly convolution neural networks (CNNs) [ 3 ], have shown significant progress in OD tasks. Deep learning-based object detectors are classified into single-shot detectors such as You Only Look Once (YOLO) and SSD (Single-Shot Detector) and object detectors based on region proposals such as RCNN, fast RCNN, faster RCNN, RFCN, and mask RCNN [ 4 ]. OD in aerial drone imagery is far more challenging as multiple factors such as altitude, camera angle, object scale, overlap, occlusion, motion blur, lack of labelled data, flat and small view of objects, real-time limited view of computation, and lack of contextual information hinder the overall object detection capabilities [ 5 ]. The tradeoff in detection accuracy and real-time performance coexists for detecting small-scale objects [ 6 , 7 , 8 ]. Single-pixel shifts cause significant interference and miss-detection due to a lack of background and foreground information [ 9 ]. These are some of the pertinent challenges for OD in UAVs [ 9 ]. Although recent approaches have significantly improved object detection performance and outcomes, the OD models are still considered opaque black boxes. It is not clear to a broader audience why the model predicts what it predicts, raising serious concerns and the need to build models that are more transparent and understandable to humans [ 10 ]. Owing to the persistent accuracy-interpretability tradeoff, i.e., the higher the complexity of the model, the lesser the interpretability, deep learning models are often viewed as black boxes. The black box nature makes understanding and comprehending their underlying working behaviour and decision-making process [ 11 ] challenging. Explainable AI (XAI) is a research field that aims to understand and interpret the working of machine learning models. It allows the interpretation of the AI-generated insights [ 12 ]. This field has gained significant attention as the use of ML models in real-world settings has increased, especially in high-risk domains such as healthcare, autonomous driving, drone-based surveillance, rescue operations [ 13 ], etc. Trustworthy and human-interpretable explanations are crucial for informed decision-making to validate the AI-generated insights in human-AI collaborative tasks considering the mission-critical application of drone imagery [ 14 ]. Drone images are subject to ethical considerations such as surveillance, security, and compliance with various data protection, privacy and transparency laws worldwide. Explainability is particularly important in critical applications such as defence, healthcare, and autonomous vehicles, where it is crucial to develop trust, transparency, and safety in machine learning models for real-world adoption and deployment [ 15 ].