Hybrid YOLO-ViT model with digital maps for efficient small traffic sign detection in real time road scenarios
R Suresha, Namrata Manohar, M Priyanka, B R Pushpa · MethodsX · 2025
Real-time tiny traffic sign identification requires high-resolution images to avoid false detection caused by the loss of contextual features, and it is computationally expensive. To overcome these challenges, the study presents a novel hybrid YOLO-ViT framework that localizes the search space and reduces computational burden by leveraging a digital map and a YOLO-ViT model. It uses a map and precise vehicle positioning, estimated by a localization module and a novel hybrid architecture that combines a YOLOv11n with Vision Transformers (ViT), enhancing detection performance, particularly in challenging conditions involving tiny, difficult-to-detect traffic signs. By integrating a map, the projection model can identify a region of interest within an image, reducing computational costs and false detections. The YOLO-ViT model combines a YOLOv11n backbone that extracts local features and a ViT to capture global contextual features. The experimental results of the proposed model achieve 97.5 % precision, 96.4 % recall, 96.1 % mAP@50, 89.62 % mAP@55:95, and 9.52 ms (≈12 fps) average inference time on the Indian road traffic sign dataset.•Design a small traffic sign detection (TSD) framework with real-time processing in challenging Indian road environments, reducing computational cost by localizing the search area and minimizing false detection through map integration.•Develop a novel hybrid YOLO-ViT model that combines the local and global contextual features to improve the accuracy of tiny TSD in challenging scenarios.•Evaluate the efficiency of the proposed model on the Indian traffic sign dataset, the Tsinghua-Tencent 100K (TT100k), and the German Traffic Sign Detection Benchmark (GTSDB), and show improvement over state-of-the-art models.