Dual-Transformer Networks with Attention-guided Fusion for Event Recognition in Aerial Videos
Javed Imran, Bajrangi Kumar Mishra, Satakshi Verma · Procedia Computer Science · 2026
Action recognition in aerial videos presents unique challenges due to varying viewpoints, scale changes, and complex backgrounds captured through Unmanned Aerial Vehicles (UAVs). This study proposes a novel multi-stream framework that leverages complementary transformer-based architectures for robust feature extraction from UAV videos. Specifically, the Inception Transformer (iFormer) and Multiscale Vision Transformer (MViT) backbones are employed to capture both local and global spatiotemporal dependencies from aerial videos. The extracted representations from each backbone are independently processed by Multi-Head Self-Attention (MHSA) layers, followed by fully connected (FC) and Softmax classification layers. To integrate the decision outputs from the two streams, a score-level fusion strategy based on the product rule is adopted, effectively emphasizing consensus predictions while reducing modality-specific uncertainties. The originality of this work lies in combining heterogeneous transformer backbones with an attention-guided decision fusion mechanism specifically designed for UAV-based action recognition. Experimental results on the ERA dataset demonstrate improved recognition performance (73.1%) compared to single-stream baselines, highlighting the benefits of combining diverse transformer architectures and attention-guided decision fusion.