CFT-YOLOv8: investigation of cross-modal fusion architectures for small-target human detection in UAV search-and-rescue

黄乐程 Huang Lecheng · DR-NTU (Nanyang Technological University) · 2026

Unmanned aerial vehicles become the core equipment in search and rescue missions, but in the most critical operating environment, the image data obtained by drones is difficult to use effectively: A visible-light camera fails in low light; an infrared camera loses its target against thermal clutter. Fusing the two modalities is the established remedy. Aerial search, however, imposes a second demand at the same time: seen from altitude, a person spans under twenty pixels, and the small-target techniques built to handle such objects have rarely been combined with cross-modal fusion. This dissertation works at that intersection. We migrate the Cross-Modality Fusion Transformer (CFT) onto the anchor-free YOLOv8 detector, rebuilding the dual-stream forward path that a single-stream framework does not provide, and confirm that the fusion mechanism survives the move: on the ground-level benchmark LLVIP it reaches the working point of the original reference (mAP50 = 0.973). On the aerial benchmark VTSaR, the same detector meets a ceiling that is easily misread as under-optimization. The baseline finds almost every target at a loose threshold (mAP50 = 0.959) yet localizes it poorly under a strict one (mAP50-95 = 0.4498). We probe this gap with five methodologically independent interventions: a higher-resolution head, a scale-aware fusion module, an ODE-inspired residual block, and two small-target regression losses. Although they disagree on where the limitation lies, all five leave mAP50-95 inside a narrow band from 0.4444 to 0.4567; the few effects that reach statistical significance run to only a few thousandths of a point. This convergence into so narrow a band is the central finding, and its cause appears structural rather than a matter of optimization. Two constraints, both read directly from the architecture, bound the accuracy beneath any single module. Cross-modal attention operates on an over-pooled representation, in which a target smaller than one pooled token is averaged into its background before a single attention weight is computed; and the label assigner routes 96.55% of positive samples to the one head that fuses by plain addition, leaving the attention-bearing layers almost unsupervised. Fusion capacity and gradient supply thus occupy disjoint levels of the feature pyramid. A third candidate, modal misalignment, is measured and set aside: the residual cross-modal offset is far too small to bind. The ceiling therefore appears structural, and this diagnosis does more than explain a shortfall. It yields a falsifiable prediction: only an intervention that closes the disjunction between fusion and gradient can raise the ceiling, and enriching either side alone cannot. Within the bounds of this architecture and this small-target regime, the contribution of this work is less a module that advances the state of the art than an account of why no such module readily does, and of the one direction in which a working one might be built.

Read the paper · More papers on PaperTik